License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmuwux5uv009b08o5bjtelh4s
Release Date: 10/6/2026
Format: WAV, TSV
Size: 1.76 GB
This dataset contains all 5,822 utterances (7.0 hours) of OpenSLR 41, "High quality TTS data for Javanese": high-quality read speech from multiple speakers, with the original audio and transcripts. It was collected by Google in collaboration with Universitas Gadjah Mada (CC BY-SA 4.0). It is re-hosted here because it was used to train an Indonesian–Javanese ASR submission to the Lost in Transcription challenge (MDC / DrivenData, 2026), whose rules require external training data to be available on MDC. An added metadata file gives the normalised transcripts used in training.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
Attribution and ShareAlike, as required by CC BY-SA 4.0. Please cite: Sodimana, K., Pipatsrisawat, K., Ha, L., Jansche, M., Kjartansson, O., De Silva, P., & Sarin, S. (2018). A Step-by-Step Process for Building TTS Voices Using Open Source Data and Framework for Bangla, Javanese, Khmer, Nepali, Sinhala, and Sundanese. Proc. SLTU 2018, 66–70. Copyright notice of the original release: Copyright 2016, 2017, 2018 Google LLC ().
Forbidden Usage
None beyond the licence (CC BY-SA 4.0 does not permit additional restrictions).
Ethical Review
(i) Nature: studio-quality read Javanese speech (a multi-speaker TTS corpus) from the original OpenSLR 41 release, with real speakers' voices. Speaker IDs are the original anonymous codes. (ii) Lawful basis: CC BY-SA 4.0 permits redistribution with attribution and share-alike. (iii) Permissions: the original publishers' CC BY-SA 4.0 release. Recording and consent were handled by the original data collectors (see Sodimana et al., 2018). (iv) Safeguards: audio and transcripts are unchanged; no personal data was added; non-exclusive listing with attribution. We ask users not to attempt speaker identification or voice cloning of the speakers (a request, not a licence condition).
Intended Use
Javanese ASR and TTS research; reproducing the Lost in Transcription competition submission.
This archive re-hosts, unchanged, the OpenSLR 41 corpus that was used to train the Indonesian–Javanese submission of participant
inyongkhafid in the Lost in Transcription challenge (Mozilla Data Collective / DrivenData, 2026). The competition requires external
training data that is not already on MDC to be published on MDC; this upload fulfils that rule.
Original resource: High quality TTS data for Javanese (SLR41), https://www.openslr.org/41/ — "Multi-speaker TTS data for Javanese (jv-ID)", developed by Google in partnership with Universitas Gadjah Mada. License: Attribution-ShareAlike 4.0 (CC BY-SA 4.0).
| Path | What |
|---|---|
jv_id_female/, jv_id_male/ | Original release folders: wavs/*.wav and line_index.tsv, unchanged (5,822 WAV files, 7.0 h). |
metadata_as_used.tsv | The 5,822 utterances used for training: id, subset, transcript (original), transcript_as_used (training normalisation). |
LICENSE | CC BY-SA 4.0 legal code. |
SHA256SUMS | SHA-256 of every file in this archive. |
transcript_as_used): spelled letters x_letter → capital letters (merged, e.g. p_letter t_letter → PT);
part-of-speech disambiguation tags (_nou, _vrb, …) removed; Unicode NFC and diacritics folded; dh → d, th → t. The audio was not modified.Released under CC BY-SA 4.0, the licence of the original corpus. Changes made: one added metadata file (metadata_as_used.tsv).
No other terms are added: CC BY-SA 4.0 does not permit additional restrictions.
Please cite the original corpus:
@inproceedings{kjartansson-etal-tts-sltu2018,
title = {{A Step-by-Step Process for Building TTS Voices Using Open Source Data and Framework for Bangla, Javanese, Khmer, Nepali, Sinhala, and Sundanese}},
author = {Keshan Sodimana and Knot Pipatsrisawat and Linne Ha and Martin Jansche and Oddur Kjartansson and Pasindu De Silva and Supheakmungkol Sarin},
booktitle = {Proc. The 6th Intl. Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU)},
year = {2018}, address = {Gurugram, India}, month = {August}, pages = {66--70}
}
jvf_04982_00083400049 — kara iki isine cuma wong lima thok wedok kabehjvf_07638_00299821958 — pabrikipun nabisco sakmenika mboten tebih saking griyane bulik prijvm_00027_00550619562 — montor toyota iku gawe gaweane entheng lan énakjvm_01519_01310118341 — sapa c_letter-en e_letter-en o_letter ne repsol pas tahun rong ewujvm_04285_00955421578 — sondre lerche gawe rèng rèngan sakdurunge ujian sosiologi dimulaijvm_08178_01897618232 — phi collins ngombe obat sakwise mangan téla karo pohungOriginal source: https://www.openslr.org/41/ (the corpus is also distributed there). Non-exclusive.
Legal basis: OpenSLR 41 was collected by Google in collaboration with Universitas Gadjah Mada and released by Google under the Creative Commons Attribution-ShareAlike 4.0 licence (Copyright 2016, 2017, 2018 Google LLC; LICENSE file at https://www.openslr.org/41/). Section 2(a)(1) of this licence lets anyone reproduce and share the material, in whole or in part, including for commercial purposes, on the conditions of attribution (Section 3(a)) and ShareAlike (Section 3(b)). This upload meets both: it keeps the same licence, credits the original publisher (copyright notice, citation and link), and contains all 5,822 utterances unchanged. Other platforms already redistribute this corpus under the same licence; for example, it is included in the "openslr" dataset on Hugging Face, which was not published by us. On 6 October 2026, at the request of the MDC reviewers, we notified the original authors (Google) that this copy is published on MDC under the same licence.