Release Date: 8/27/2026
Format: MP3, TSV
Size: 245.00 MB
Share
The Khowar ASR Dataset is a resource designed to support research in speech and language technologies for the Khowar language. It contains utterance-level audios and corresponding transcriptions, extracted from ELAN files. This dataset aims to advance the development of AI and NLP technologies for the underrepresented Khowar language while supporting language documentation, preservation, and future multilingual research.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($100.00) platform fee paid to MDC.
Licensing
MDC Data License Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, voice-print matching, or any other technique intended to link recordings to real individuals.
Forbidden Usage
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, voice-print matching, or any other technique intended to link recordings to real individuals.
Khowar (ISO 639-3 code khw) is a Dardic language spoken primarily in Chitral district in Khyber Pakhtunkhwa, Pakistan, with additional speaker communities in Gilgit-Baltistan and among diaspora in cities such as Peshawar and Karachi. It has a well-developed literary tradition relative to many other Dardic languages, including active poetry, folklore, and language-preservation organizations, yet it remains under-resourced for speech and NLP technologies. This dataset includes a notable number of clips from speakers directly engaged in Khowar language and literature preservation work.
The dataset comprises unscripted, first-person interviews and talks in Khowar featuring speakers from Chitral and surrounding communities. Topics span personal and educational journeys (life in primary school, university life, higher education, M.Phil research, admission in law college, hostel life), professional life (career shifts such as moving to Peshawar for a new job, office work experiences, work for Khowar language and literature preservation), cultural heritage (the Pathak Festival, Khowar folktales including a recurring Fairy of the Sky and Eagle tale series, wedding customs of Chitral, traditional dishes such as tarbrad, village life in Bumborate and Kosht), social and environmental topics (proper use of water, mitigation of environmental disasters, heat waves), and closing reflections and community messages (favourite hobbies and sports, cherished childhood memories, the Chitral public library, and awards for language development).
The dataset consists of a clips/ directory with 10,000 segmented audio clips, taken from the original 152 videos transcribed in ELAN, and a data.tsv file, which contains three columns:
video_id: an identifier of the original video the clip was extracted from. This is also useful for lookup in the metadata file, discussed below.
path: the relative path to the audio clip (e.g. clips/.mp3
sentence: the Khowar transcript.
In addition, a separate TSV metadata file accompanies the dataset, providing speaker-level and clip-level metadata for every original recording, including:
video_id (original ID of the recording corresponding to the folder in the multimodal dataset)
duration
description
speaker_count
speaker_gender
speaker_age_group