License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmthb06nd039ynw07uq3pahj7
Release Date: 8/31/2026
Format: WAV, WEBM, TSV
Size: 1.12 GB
Share
This dataset contains text-to-speech recordings from 16 English language learners with various social backgrounds and dialects among Indonesian speakers.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is restricted to non-commercial use and is intended exclusively for research, machine learning development, and internal use. The internal use means the dataset is not allowed to be re-hosted, re-shared, and re-distributed so the dataset access is granted only from the MDC platform. Any use beyond these purposes, including commercial applications, requires explicit authorization and is not permitted under the current usage terms.
Forbidden Usage
The dataset must not be used for voice cloning, speaker imitation, or the development of models designed to replicate the identities, voices, or speech patterns of the specific individuals represented in the corpus. Using this dataset to train chatbots, large language models, or other AI systems that imitate specific speakers is strictly prohibited. Additionally, the dataset is not permitted for commercial purposes.
Ethical Review
All participants were informed and gave consent to make this dataset. A few participants identify their English language levels by taking pre-test. Participants used their own device and read individually the provided text in the project corpus at https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/sabre. Each participant read 1000 sentences between 1 and 10 words long. Finally, the collection of audio recordings was compiled into a comprehensive dataset.
Intended Use
This dataset is intended to support research, AI development, education, and data science initiatives. It may be used for improving AI systems, developing machine learning models, and advancing research in computational and educational applications. Its primary purpose is to facilitate innovation, experimentation, and learning within academic and development-oriented environments.
English
There are 16 speakers from different ages between 20 and 40, sociocultural backgrounds, and proficiency levels. There are 6 advanced learners, 9 speakers with intermediate levels as well as 1 speaker with beginner level. Most of them are native bilingual Indonesian and local languages of Indonesia (Javanese, Sundanese, Mandarnese).
The source text consists of 1,001 sentences taken from this multilingual readability corpus. The sentences are from OpenSubtitles, and are between 1 and 10 words long.
Approximately 11.5 hours
Public open access is permitted with proper attribution and citation of the dataset source.
Abdul Cholik Al Aziz, Adiguna Suwondo, Alifa Ardla Farahdila, Dhila Damayanti, Diny Opticawati, Masyhuri Farhan, Feby Nur Dianingtyas, Muqsit Harjuno, Lailatul Zunaeva, Puput Rizky Afriani, Rahmah Nabilatun Nisa’, Sinta Dwi Kurnia, Yunus Aji Lumintang. (2026). Indonesian-Learner-English [Data set]. Mozilla Data Collective. URL [dataset link].
The dataset contains a TSV file, metadata.tsv, with the following columns:
audio_id: A unique identifier for each audio recording, formatted as speaker_id-recording_id.
speaker_id: A unique identifier assigned to each speaker.
audio_filename: The filename of the corresponding audio recording.
sentence: The text corresponding to the audio recording.
num_attempts: The number of recording attempts made by the speaker. Speakers were asked to read each sentence as fluently as possible and were encouraged to retake the recording if they encountered difficulties during the reading.
Compensated · similar languages
A Khowar-English speech translation dataset with approximately 9.5 hours of Khowar audio with English text translations.
A Khowar multimodal dataset with 10 hours of synchronized video, audio, transcriptions, speech segmentation, and ELAN annotations.