License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmtr89f6505ilo107jwdqu6zg
Release Date: 9/7/2026
Format: WAV, WEBM, TSV
Size: 1.66 GB
Share
This dataset contains text-to-speech recordings from 17 English language learners with various social backgrounds and dialects among Indonesian speakers.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
The dataset is provided exclusively for scientific, academic, and educational purposes. For commercial use of derivative products or other integration into commercial production systems, we kindly ask you to notify us through a brief and clear introduction on MDC platform and please write an email to: [email protected], [email protected], [email protected], [email protected], [email protected], [email protected]
Forbidden Usage
The dataset must not be used for voice identification or voice-type classification (deanonymize/re-identify any speaker in the dataset). It must also not be used to generate, distribute, or facilitate the spread of false information, misinformation, disinformation, or other forms of manipulative content.
Ethical Review
All participants were informed and gave consent to make this dataset. A few participants identify their English language levels by taking pre-test. Participants used their own device and read individually the provided text in the project corpus at https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/sabre. Each participant read 1000 sentences between 1 and 10 words long. Finally, the collection of audio recordings was compiled into a comprehensive dataset.
Intended Use
The dataset is intended for research, artificial intelligence (AI) development, and educational purposes; including language research, AI-based applications, and the development of educational platforms.
There are 17 speakers from different ages between 20 and 30, sociocultural backgrounds, and proficiency levels. There are 6 speakers with advanced and 11 and intermediate level. Most of whom are native bilingual Indonesian and local languages of Indonesia (Bugis, Javanese, Mandar, and Sundanese).
The source text consists of 1,001 sentences taken from (https://github.com/aconeil/Readability) multilingual readability corpus. The sentences are from OpenSubtitles, and are between 1 and 10 words long.
Approximately 12 hours (17 files)
This dataset is open to the public with proper source attribution or citation. In addition, the dataset owner accepts compensation in any form either for commercial or non-commercial use.
Adiguna Suwondo, Dhila Damayanti, Diny Opticawati, Jian Alina, Masyhuri Farhan, Muhammad Hasrinur Ridho, Putri Sholikha, Rosyid Hidayatul Fadilah, Sinta Dwi Kurnia, Surti Syafiurahmi, Yunus Aji Lumintang. (2026). IndoLearner-English Speech Corpus [Dataset]. Mozilla Data Collective. URL [dataset link]
The dataset contains a TSV file, metadata.tsv, with the following columns:
audio_id: A unique identifier for each audio recording, formatted as speaker_id-recording_id.
speaker_id: A unique identifier assigned to each speaker.
audio_filename: The filename of the corresponding audio recording.
sentence: The text corresponding to the audio recording.
num_attempts: The number of recording attempts made by the speaker. Speakers were asked to read each sentence as fluently as possible and were encouraged to retake the recording if they encountered difficulties during the reading.
Compensated · similar languages
100 hours of premium, rights-cleared English-language scripted film content, with English subtitles and accompanying metadata, available for AI training
100k native-reviewed colloquial sentences translated from US English into Yoruba for fine-tuning conversational MT and AI models.