Release Date: 9/1/2026
Format: MP3, TSV
Size: 16.24 MB
Share
This dataset contains 35 minutes of transcribed conversational speech from Spanish-English bilinguals living in North America. It is intended to be used for post-training, benchmarking, development, and/or tuning ASR models for the "Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts" competition.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
You agree not to attempt to determine the identity of speakers in this dataset. Any attempt to clone the voice or train models that imitate the speakers in this dataset is forbidden.
Forbidden Usage
You agree not to attempt to determine the identity of speakers in this dataset. Any attempt to clone the voice or train models that imitate the speakers in this dataset is forbidden.
Intended Use
Post-training, parameter-tuning, evaluation, cross-validation, etc. of ASR models.
The development dataset for the Spanish-English track of the "Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts" competition from Mozilla Data Collective. More information about the competition can be found at the competition website.
Spanish and English are the two most widely-spoken languages in the Unites States, and there are high rates of Spanish-English bilingualism in North America. Nearly 40% of U.S. Hispanics are bilingual in English and Spanish, and Spanish is also used by over 2 million non-Hispanics in the U.S. (https://www.pewresearch.org/short-reads/2015/03/24/a-majority-of-english-speaking-hispanics-in-the-u-s-are-bilingual/). Code-switching between Spanish and English among these populations is commonplace.
The speech data was recorded by speakers via WhatsApp voice notes. These were first automatically transcribed using omniASR_LLM_3B, and manually post-edited by the participants. They were also annotated for turn-level language and for PII. PII-containing turns were omitted.
The dataset consists of 35 minutes of transcribed speech from 2 speakers. Most turns contain both Spanish and English, with a small number containing only Spanish.
./metadata.tsv
./clips/
.mp3
audio_filename: The name of the file
speaker: A speaker ID
transcript: The transcript
language: en, spa, or enspa
convo_id: The conversation ID
Compensated · similar languages
A Khowar-English speech translation dataset with approximately 9.5 hours of Khowar audio with English text translations.
A Khowar multimodal dataset with 10 hours of synchronized video, audio, transcriptions, speech segmentation, and ELAN annotations.