License:
CC-BY-SA-4.0
Steward:
Bahasa BahanaDataset ID:
cmrxfn05r014znu07ub3muel9
Release Date: 7/23/2026
Format: WEBM, WAV, TSV
Size: 1.14 GB
Share
This dataset contains read speech recordings from 14 English language learners of Indonesia, speakers with across various language levels and social backgrounds .
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
This dataset is intended for NLP (AI training), teaching, and research. However, if you wish to use it for other purposes, please contact the administrator by submitting an access request with a clear statement of your intended use. Use of the dataset for commercial derivative products requires contacting the admin, and Bahasa Bahana will receive any type of compensation, which will be discussed further. This dataset is open to the public and is not exclusively granted to any party. This dataset is released for open use. It may be used for research, education, and commercial applications by a company earning not over 10 million dollars per year. For commercial use, the data user must send explicit permission through email to the data owner. For derivative products or modification, users must explain explicitly the modification and respect privacy, copy right, and law. Users must comply with the dataset license terms and provide appropriate attribution for research uses, article publications, and products.
Forbidden Usage
You must not attempt to identify or re-identify any individual speaker. You must not use this dataset to clone voices or create systems that imitate specific speakers. You must not use this dataset for malicious, deceptive, or harmful purposes according to laws and social norms. You must not use this dataset to generate misleading, hate-speech, discrimination, bullying, and fraudulent audio content. Re-distribute the dataset into an individual or organization is not allowed. Accessing this dataset is only via Mozilla Data Collective and comply with the terms and conditions. Any use that is intended to commit war and violates privacy rights, human rights (animate and inanimate), or applicable laws is strictly prohibited. This dataset may not be used by companies, individuals, or affiliates with any known human rights violations or war crimes. Unattributed use of the dataset and causing disadvantages to contributors is prohibited.
Ethical Review
All participants were informed and gave consent to make this dataset. A few participants identify their English language levels by taking pre-test. Participants used their own device and read individually the provided text in the project corpus at https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/sabre. Each participant read 1000 sentences between 1 and 10 words long. Finally, the collection of audio recordings was compiled into a comprehensive dataset.
Intended Use
This dataset is made for development of NLP (AI training), language teaching-learning, ASR for English learners, language learning applications, computer-aided language learning, and linguistics research for second language learning in Indonesia.
English
There are 14 speakers from different ages, social backgrounds, and professions. There are 1 speaker with beginner, 10 intermediate, and 3 advanced level. All of them are native bilingual Indonesian and regional languages (Manggarai, Javanese, Betawi) of Indonesia. Inside the dataset, the initial code of folder naming indicates the L1 of the speakers, for example manggarai_beginner, jav_intermediate, and jav_advanced. Their ages are in their 20s and 30s.
The source text consists of 1,001 sentences taken from this multilingual readability corpus (https://github.com/aconeil/Readability). The sentences are from OpenSubtitles, and are between 1 and 10 words long.
10 hours
This dataset is open to the public with proper source attribution or citation. In addition, the dataset owner accepts compensation in any form either for commercial or non-commercial use of this dataset.
Jian Alina, Adit Bondan Pradhana, Lailatul Zunaeva, Nur Khafidzah, Linda Rahmawati, Muhammad Hasrinur Ridho. (2026). BAHANA-Speech Corpus of English Learners from Indonesia [Data set]. Mozilla Data Collective. URL [dataset link].
The dataset contains a tsv file, metadata.tsv, with the following columns:
audio_id: a key with speaker_id-audio_id
speaker_id
audio_filename
sentence: text
num attempts: Speakers were asked to read the sentence as fluidly as possible, and encouraged to do retakes if they struggled during a reading. This column shows how many attempts were taken to record the sentence.