Release Date: 9/25/2026
Format: MP3, TSV
Size: 1.24 GB
A Tagalog speech recognition dataset consisting of 45+ hours of transcribed conversations recorded via voice memos on WhatsApp and Telegram, from around 50 speakers.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($450.00) platform fee paid to MDC.
Licensing
MDC Data License Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, voice-print matching, or any other technique intended to link recordings to real individuals.
Forbidden Usage
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, voice-print matching, or any other technique intended to link recordings to real individuals.
Intended Use
ASR training and evaluation
This dataset consists of over 45 hours of spoken conversation in Tagalog from 48 distinct speakers.
Participants engaged in two-party conversations via chat applications (WhatsApp, Telegram), communicating exclusively via voice notes. The audio from these conversations was automatically transcribed using the omniASR-LLM-1B Omnilingual ASR model. The automated transcriptions were then post-edited to create the gold transcriptions released in the dataset.
The conversations do not follow any strict topics or themes, as speakers were instructed to converse naturally. There are a number of instances of English code-switching in the dataset. There is approximately 1 hour of audio per speaker.
Some basic information about the
47 hours of audio
~3,000 sentences
Avg 1 minute per turn
~300,000 words (~30,000 unique words)
The dataset consists of a clips/ directory containing all of the utterance audio MP3 files, and a mapping.tsv file with the transcriptions, file names, and other metadata. The columns are:
audio_filename: the name of the audio file (mp3) contained in the clips/ directory
id: an identifier for the utterance (the filename without the extension)
speaker: a number corresponding to the speaker
conversation_id: an identifier for the conversation the turn is a part of (where available)
transcript: The reviewed transcription
offensive: A binary label for whether the transcription/audio contains potentially offensive content,
duration: length of the audio clip in seconds