License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmuwuxgta009d07mfzupkk4b0
Release Date: 10/6/2026
Format: FLAC, TSV
Size: 713.26 MB
This dataset contains 8,000 read Javanese utterances (12.7 hours, 786 speakers) with their original audio and transcripts, plus the full transcript file of the corpus (185,076 lines). It is a subset of OpenSLR 35, the "Large Javanese ASR training data set" collected by Google in collaboration with Reykjavik University and Universitas Gadjah Mada (CC BY-SA 4.0). The subset was used to train an Indonesian–Javanese ASR submission to the Lost in Transcription challenge (MDC / DrivenData, 2026), whose rules require external training data to be available on MDC. An added column gives the normalised transcripts used in training.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
Attribution and ShareAlike, as required by CC BY-SA 4.0. Please cite: Kjartansson, O., Sarin, S., Pipatsrisawat, K., Jansche, M., & Ha, L. (2018). . Proc. SLTU 2018, 52–55.
Forbidden Usage
None beyond the licence (CC BY-SA 4.0 does not permit additional restrictions).
Ethical Review
(i) Nature: read Javanese speech from the original OpenSLR 35 release, with real speakers' voices. Speakers are identified only by the anonymous IDs of the original release; there are no names or demographic data. (ii) Lawful basis: CC BY-SA 4.0 permits redistribution with attribution and share-alike. (iii) Permissions: the original publishers' CC BY-SA 4.0 release. Recording and consent were handled by the original data collectors (see Kjartansson et al., 2018). No new personal data was added. (iv) Safeguards: audio and transcripts are unchanged; no personal data was added; non-exclusive listing with attribution. We ask users not to attempt speaker identification (a request, not a licence condition).
Intended Use
Training and evaluating Javanese ASR; reproducing the Lost in Transcription competition submission.
This archive re-hosts, unchanged, the part of the OpenSLR 35 corpus that was used to train the Indonesian–Javanese
submission of participant inyongkhafid in the Lost in Transcription challenge (Mozilla Data Collective / DrivenData, 2026).
The competition requires external training data that is not already on MDC to be published on MDC; this upload fulfils that rule.
Original resource: Large Javanese ASR training data set (SLR35), https://www.openslr.org/35/ — "Javanese ASR training data set containing ~185K utterances", developed by Google in partnership with Reykjavik University and Universitas Gadjah Mada. License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0).
| Path | What |
|---|---|
audio/{xx}/{id}.flac | 8,000 utterances, original FLAC files from the OpenSLR release, not re-encoded (12.7 h, 786 speakers). A fixed random subset (seed 20260827) of the full corpus. |
utt_spk_text.tsv | Original transcript index of the full corpus (185,076 lines: FileID, UserID, transcription), unchanged. All of these transcripts (text only, no audio) were used for text-only adaptation of one model. |
subset_used.tsv | The audio utterances used: id, speaker, transcript (original), transcript_as_used (training normalisation, see below). |
LICENSE | CC BY-SA 4.0 legal code. |
SHA256SUMS | SHA-256 of every file in this archive. |
transcript_as_used): Unicode NFC; diacritics folded (è/é/ê/ě → e, ō → o, ā → a); typographic quotes → ASCII;
bracketed tags removed; dh → d, th → t (to match the competition's Javanese spelling); whitespace collapsed. The audio was not modified.Released under CC BY-SA 4.0, the licence of the original corpus. Changes made: selection of a subset, and one added column (transcript_as_used).
No other terms are added: CC BY-SA 4.0 does not permit additional restrictions.
Please cite the original corpus:
@inproceedings{kjartansson-etal-sltu2018,
title = {{Crowd-Sourced Speech Corpora for Javanese, Sundanese, Sinhala, Nepali, and Bangladeshi Bengali}},
author = {Oddur Kjartansson and Supheakmungkol Sarin and Knot Pipatsrisawat and Martin Jansche and Linne Ha},
booktitle = {Proc. The 6th Intl. Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU)},
year = {2018}, address = {Gurugram, India}, month = aug, pages = {52--55}
}
09fc6cb1a1 — wetone Nur kuwi Setu Pon0d867c3a26 — Victoria Beckham seneng ngrungokne musik rap12a9f95411 — Any insect or spider may often be called a bug593428d698 — Henry banjur dadi tentara lan oleh tugas pitung taunccc43ffe78 — Susu saka sawetara kéwan dikumpulaké manungsa kanggo diombéed7a93ff93 — Lance Armstrong kuwi alumni Stamford UniversityOriginal source: https://www.openslr.org/35/ (the corpus is also distributed there). Non-exclusive.
Legal basis: Google released OpenSLR 35 under the Creative Commons Attribution-ShareAlike 4.0 International licence (LICENSE file at https://www.openslr.org/35/). Section 2(a)(1) of this licence lets anyone reproduce and share the material, in whole or in part, including for commercial purposes, on two conditions: attribution (Section 3(a)) and ShareAlike (Section 3(b)). This upload meets both: it keeps the same licence, credits the original publisher (copyright notice, citation and link), and its audio and transcripts are unchanged (8,000 of the ~185,000 original utterances). Other platforms already redistribute this corpus under the same licence; for example, it is included in the "openslr" dataset on Hugging Face, which was not published by us. On 6 October 2026, at the request of the MDC reviewers, we notified the original authors (Google) that this copy is published on MDC under the same licence.