Release Date: 8/28/2026
Format: TSV
Size: 407.20 KB
Share
A parallel corpus of over 10,000 sentences of Khowar and English. The sentences are derived from a multimodal corpus consisting of 152 transcribed and translated Khowar language videos.
Licensing
MDC Data License Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
Any attempt to identify, re-identify, or expose the identity of speakers of the texts featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, or any other technique intended to link texts to real individuals.
Forbidden Usage
Any attempt to identify, re-identify, or expose the identity of speakers of the texts featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, or any other technique intended to link texts to real individuals.
Khowar (ISO 639-3 code khw) is a Dardic language spoken primarily in Chitral district in Khyber Pakhtunkhwa, Pakistan, with additional speaker communities in Gilgit-Baltistan and among diaspora in cities such as Peshawar and Karachi. It has a well-developed literary tradition relative to many other Dardic languages, including active poetry, folklore, and language-preservation organizations, yet it remains under-resourced for speech and NLP technologies. This dataset includes a notable number of clips from speakers directly engaged in Khowar language and literature preservation work.
The dataset comprises transcripts and translations from unscripted, first-person interviews and talks in Khowar featuring speakers from Chitral and surrounding communities. Topics span personal and educational journeys (life in primary school, university life, higher education, M.Phil research, admission in law college, hostel life), professional life (career shifts such as moving to Peshawar for a new job, office work experiences, work for Khowar language and literature preservation), cultural heritage (the Pathak Festival, Khowar folktales including a recurring Fairy of the Sky and Eagle tale series, wedding customs of Chitral, traditional dishes such as tarbrad, village life in Bumborate and Kosht), social and environmental topics (proper use of water, mitigation of environmental disasters, heat waves), and closing reflections and community messages (favourite hobbies and sports, cherished childhood memories, the Chitral public library, and awards for language development).
The dataset consists of a TSV file with one column containing the Khowar text and the other containing the English text.
Compensated · similar languages
A Khowar-English speech translation dataset with approximately 9.5 hours of Khowar audio with English text translations.
A Khowar ASR dataset with approximately 10 hours of audio with transcriptions.