Release Date: 8/26/2026
Format: MP3, TSV
Size: 35.36 MB
Share
This dataset contains 1 hour and 18 minutes of transcribed conversational speech from Nahuatl-Spanish bilinguals, from three different Nahuatl varieties (Highland Puebla Nahuatl, Western Sierra Puebla Nahuatl, and Western Huasteca Nahuatl). It is intended to be used for post-training, development, and/or tuning ASR models for the "Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts" competition.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
You agree not to attempt to determine the identity of speakers in this dataset. Any attempt to clone the voice or train models that imitate the speakers in this dataset is forbidden.
Forbidden Usage
You agree not to attempt to determine the identity of speakers in this dataset. Any attempt to clone the voice or train models that imitate the speakers in this dataset is forbidden.
Intended Use
Post-training, parameter-tuning, evaluation, cross-validation, etc. of ASR models.
The development dataset for the "Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts" competition from Mozilla Data Collective. More information about the competition can be found at the competition website.
Nahuatl is spoken by over 1.5 Million people in Mexico, and is made up of a number (~30) of distinct linguistic varieties. The variants included in this dataset are Highland Puebla Nahuatl (ISO 639-3 azz), Western Sierra Puebla Nahuatl (ISO 639-3 nhi) and Western Huasteca Nahuatl as spoken in San Luis Potosí (ISO 639-3 nhw).
Many Nahuatl speakers today are bilingual in Spanish. As a result, and as a result of 5 centuries of close language contact with Spanish, everyday Nahuatl speech frequently contains many loanwords, calques, and code-switching.
The speech data was recorded by speakers via WhatsApp voice notes. These were first automatically transcribed using omniASR_LLM_3B, and manually post-edited by the participants. They were also annotated for turn-level language and for PII. PII-containing turns were omitted.
The dataset consists of 1 hour and 18 minutes of transcribed speech, with approximately between 20 and 30 minutes of each variant.
The transcriptions use a popular orthography that uses "k" for /k/, "s" for /s/, "u" for /w/, and "j" for /h/. Spanish words that have not been phonologically adapted are written using the Spanish conventions.
./metadata.tsv
./clips/
.mp3
audio_filename: The name of the file
speaker: A speaker ID
transcript: The transcript
language: azz, nhi, or nhw
convo_id: The conversation ID