License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmuzv3veu00a207nxe6lb3gxd
Release Date: 10/8/2026
Format: WAV, WEBM, TSV
Size: 1.26 GB
This dataset contains read speech recordings from 14 Spanish language learners of Indonesia.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
The dataset is primarily intended for scientific research, academic activities, and educational purposes. Commercial use, including direct commercialization or integration into commercial production systems, is allowed with prior notification to the dataset owner. Users should briefly state their intended use when submitting an access request. Users must comply with the applicable licensing requirements and acknowledge and respect the rights and conditions set by the dataset owner.
Forbidden Usage
The dataset must not be used for specific voice cloning, specific speaker imitation, or any attempt to replicate the identity or vocal characteristics of individuals. It is also prohibited to use the dataset for identifying speakers or analyzing speaker-specific voice attributes. Furthermore, the dataset must not be used to generate, distribute, or support the creation of misinformation, disinformation, or other forms of manipulative content. Any modification or alteration of the dataset without prior approval from the dataset owner is strictly prohibited.
Ethical Review
All participants were informed and gave consent to make this dataset.Participants used their own device and read individually the provided text in the project corpus at https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/sabre. Each participant read 1000 sentences between 1 and 10 words long. Finally, the collection of audio recordings was compiled into a comprehensive dataset.
Intended Use
The dataset is intended to support research in linguistics, particularly studies related to foreign language learning and learner language development. It may also be used for the development and evaluation of artificial intelligence systems, as well as educational applications that enhance language learning processes. The dataset is designed to support academic research, AI development, and educational purposes.
Spanish
There are 14 speakers from different ages, social backgrounds, and professions. Most of them are native bilingual Indonesian and regional languages(Javanese, Mandarese, Riau, Sundanese. Inside the dataset, the initial folder name indicates the L1 of the speakers, for example Mandar_Beginner, Jawa_Beginner, and Sunda_Beginner.
The source text consists of 1,001 sentences taken from this multilingual readability corpus. The sentences are from OpenSubtitles, and are between 1 and 10 words long.
10 hours
Public open access is permitted with proper attribution and citation of the dataset source.
Adiguna Suwondo, Alifa Ardla Farahdila, Muhammad Andrean, Aninda Ayu Sitorismi, Dhila Damayanti, Jian Alina, Muqsit Harjuno, Linda Rahmawati, Sinta Dwi Kurnia, Yunus Aji Lumintang. (2026). IDN Spanish Corpus Learner [Data set]. Mozilla Data Collective. URL [dataset link].
The dataset contains a TSV file, metadata.tsv, with the following columns:
audio_id: A unique identifier for each audio recording, formatted as speaker_id-recording_id.speaker_id: A unique identifier assigned to each speaker.audio_filename: The filename of the corresponding audio recording.sentence: The text corresponding to the audio recording.num_attempts: The number of recording attempts made by the speaker. Speakers were asked to read each sentence as fluently as possible and were encouraged to retake the recording if they encountered difficulties during the reading.Compensated · similar languages
100 hours of rights-cleared Spanish-language episodic television content (5 episodes/series), with English subtitles and structured metadata
20 hours of rights-cleared Castilian Spanish-language scripted film content, accompanied by structured metadata and English Subtitles