License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmt6ft8jp02vanz075sg7xm0m
Release Date: 8/23/2026
Format: WAV, WEBM, TSV
Size: 1.06 GB
Share
This dataset contains text-to-speech recordings from 14 English language learners with various social backgrounds and dialects among Indonesian speakers.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
The dataset is intended exclusively for research, machine learning development, and internal use. Any commercial exploitation, redistribution for commercial purposes, or use in commercial products and services is prohibited. Users must ensure that the dataset is used only within the permitted scope and in accordance with applicable ethical and usage guidelines.
Forbidden Usage
The dataset is strictly prohibited for commercial purposes. Any use of the dataset for voice cloning, speaker imitation, or the development of systems designed to replicate the identity, characteristics, or behaviour of specific speakers is not allowed. Additionally, the dataset must not be used to train chatbots, large language models, or other AI systems that attempt to simulate or impersonate the original speakers.
Ethical Review
All participants were informed and gave consent to make this dataset. Participants used their own device and read individually the provided text in the project corpus at https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/sabre. Each participant read 1000 sentences between 1 and 10 words long. Finally, the collection of audio recordings was compiled into a comprehensive dataset.
Intended Use
The dataset is intended to support research activities, data science development, AI improvement, and educational purposes. It may be used for developing and evaluating machine learning models, advancing artificial intelligence technologies, and supporting studies in language learning, computational linguistics, or related fields. The dataset is designed to contribute to academic and technological development while maintaining responsible and ethical usage.
Spanish
There are 14 contributors in total, consisting 13 beginners and 1 intermediate level, from different ages, social backgrounds, and professions. Most of them are native bilingual Indonesian and regional languages (Sundanese and Javanese) of Indonesia. Inside the dataset, the initial folder name indicates the L1 of the speakers, for example Sunda_beginner and Jawa_Beginner.
The source text consists of 1,001 sentences taken from this multilingual readability corpus. The sentences are from OpenSubtitles, and are between 1 and 10 words long.
10 hours with 14 files
Public open access is permitted with proper attribution and citation of the dataset source. Recommended citation:
Abdul Cholik Al Aziz, Diny Opticawati, Masyhuri Farhan, Feby Nur Dianingtyas, Jian Alina, Muqsit Harjuno, Nur Khafidzah, Puput Rizky Afriani, Rahmah Nabilatun N., Sinta Dwi Kurnia, Surti Syafiurahmi, and Yunus Aji Lumintang. (2026). Indonesian-Spanish Learner Corpus [Data set]. Mozilla Data Collective. URL [dataset link].
The dataset contains a TSV file, metadata.tsv, with the following columns:
audio_id: A unique identifier for each audio recording, formatted as speaker_id-recording_id.
speaker_id: A unique identifier assigned to each speaker.
audio_filename: The filename of the corresponding audio recording.
sentence: The text corresponding to the audio recording.
num_attempts: The number of recording attempts made by the speaker. Speakers were asked to read each sentence as fluently as possible and were encouraged to retake the recording if they encountered difficulties during the reading.