License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmut2t2wg02binv074906cueh
Release Date: 10/4/2026
Format: WAV, TSV
Size: 5.25 GB
13,469 fully synthetic utterances (about 61 hours) of informal Javanese, Indonesian and code-mixed speech, created to train ASR models for the Indonesian–Javanese track of the Lost in Transcription challenge (MDC / DrivenData, 2026). Texts were generated locally by Gemma 4 26B-A4B. Speech was produced by VoxCPM2 with a LoRA adapter trained on OpenSLR 41 and the MDC dataset "TTS Central Javanese", using textual persona prompts; no reference recording was used to clone a voice. No real person was recorded.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
Attribution and ShareAlike, as required by CC BY-SA 4.0. Attribution: OpenSLR 41 (Sodimana et al., 2018) and TTS Central Javanese (Fithrotunnisa, R., 2026, Mozilla Data Collective ).
cml5bn4k900aame07u0rwidcgForbidden Usage
None beyond the licence (CC BY-SA 4.0 does not permit additional restrictions).
Ethical Review
(i) Nature: fully synthetic speech and machine-generated text; no human was recorded and no original recording is included. The TTS adapter was trained on recordings of real speakers (OpenSLR 41, TTS Central Javanese), so some voices may resemble them. (ii) Lawful basis: both upstream datasets are CC BY-SA 4.0 and this adaptation is released under the same licence; the competition rules require publication on MDC. (iii) Permissions: CC BY-SA 4.0 for OpenSLR 41; written permission from the owner of TTS Central Javanese (30 September 2026) covers this use and this publication; the outputs of the text-generation model are free to use. (iv) Safeguards: voices were designed from text prompts without cloning a reference recording; the transcripts contain no personal data; the texts were checked against the competition development transcripts (0 shared 8-word sequences). We ask users not to use the voices for speaker identification or impersonation (a request, not a licence condition).
Intended Use
Training and augmenting Javanese, Indonesian and code-mixed ASR; reproducing the Lost in Transcription competition submission.
13,469 fully synthetic utterances (61.2 h) of informal Javanese, Indonesian and code-mixed Javanese–Indonesian speech. The dataset was created by
participant inyongkhafid to train ASR models for the Lost in Transcription challenge (Mozilla Data Collective / DrivenData, 2026),
and is published on MDC because the competition requires external training data to be published on MDC.
No real person was recorded for this dataset. All texts are machine-generated, and all audio is produced by a text-to-speech model.
| Path | What |
|---|---|
audio/{subset}/{id}.wav | 13,469 WAV files, 16,000 Hz mono, exactly as used for training. |
metadata.tsv | id, file, subset, language_tag (Indonesian or empty = Javanese / mixed), duration_s, transcript. |
LICENSE · NOTICE | CC BY-SA 4.0 legal code · attribution of the upstream datasets. |
SHA256SUMS | SHA-256 of every file in this archive. |
Subsets (generation batches): sintetis2 1,660 clips (8.3 h) · sintetis3 7,445 clips (35.5 h) · sintetis5 4,364 clips (17.4 h)
sintetis5) is Indonesian conversational text, partly rewritten deterministically. Overlap check against the development
transcripts: 0 of 2,611 unique texts share an 8-word sequence; 1 text (4 clips) shares one generic 6-word phrase about a price.Together with real corpora (Jember Javanese Spontaneous Speech Corpus, OpenSLR 35/41, Common Voice, MDC podcasts, TTS Central Javanese), these clips fine-tuned the members of the Indonesian–Javanese competition ensemble: Qwen3-Omni-30B-A3B, Qwen3-ASR-1.7B, omniASR-LLM-7B and Whisper large-v3 LoRA. The 4,519 original TTS Central Javanese recordings that were also used in training are not included here; they are available on MDC.
CC BY-SA 4.0. The voices derive from a model adapted on OpenSLR 41 and TTS Central Javanese, both CC BY-SA 4.0. The owner of TTS Central Javanese granted permission for this use and for publishing the synthetic data (30 September 2026). No additional terms are added.
cml5bn4k900aame07u0rwidcg.The speech is synthetic, but the TTS adapter was trained on recordings of real speakers, so some voices may resemble them. Please do not use this dataset for speaker identification, voice cloning or impersonation (a request, not a licence condition). The transcripts contain no personal data.
s2_0013_p2 — Pas lagi di pasar... terus aku nyoba beli cenil yang warnanya ijo gitu. Terus pas lagi mangan... e... teksturnya kenyal banget sampe gigiku kayak nyangkut. Aduh, susah banget... eh, maksudnya malah jadi susah nelen. Gak estetik banget deh.s2_0076_p5 — Mbiyen pas musim tanam, rame banget lho, gitu... kabeh-kabeh bareng-bareng ngolah sawah. Saiki kan... ya gitu, wong-wong wis mulai tinggal neng kota, sawah yo diurus karo sing wis tuwa-tuwa wae. Dadi koyo... sedih gitu.s3_r5f1228_1364_p1 — Akhire biji ujianne wis metu. Alhamdulillah, angkane apik-apik banget, gak elek. Anakku seneng banget, terus njaluk tuku bakso... e... bakso urat neng pasar. Yo nggo ngerayakke merga wis lulus ujian sing angel kuwi.s3_r5j1x0495_451_p4 — Dalan-dalan nang daerah Jawa Tengah kuwi... e... kadhang akèh bolongane, ya. Dadine nek liwat wengi-wengi, ngeri banget ngrasakne guncangane. Terus... ngono, stir-e dadi sering goyang-goyang dewe. Kaya... eh, maksudé setire dadi abot banget nek kena lubang ngono.s3_r5j2x0106_090_p0 — Terus si Agus crito, angin sore iki kan semilir banget... e... penak nggo miberke layangan. Dheweke ngomong, nek pengen layangane mumbul dhuwur, awakdewe kudu pinter ngulur benang... terus ojo sampek gelasane adoh banget, ben ora angel narike.s5_r5f0816_970_p0 — Aku pengen deh, gaya berpakaianku kayak artis yang lagi naik daun itu. Yang bajunya selalu kelihatan mewah. Nggak perlu yang mahal sih, yang penting, ya, yang mirip-mirip dikit lah. Biar kalau lagi jualan di pasar, penampilanku, jadi lebih segar dilihat orang.