Release Date: 8/19/2026
Format: MP4, WAV, EAF
Size: 167.23 GB
Share
The Khowar Multimodal Dataset is a comprehensive language resource designed to support research in speech and language technologies for the Khowar language. It contains synchronized video and audio recordings accompanied by time-aligned transcriptions, speech segmentation, and ELAN annotation files, enabling fine-grained multimodal analysis. The dataset is suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), speech segmentation, forced alignment, speaker and gesture analysis, multimodal language understanding, and linguistic research. By combining audiovisual data with rich annotations, it provides valuable resources for studying the interaction between speech, facial expressions, and other non-verbal communication cues. This dataset aims to advance the development of AI and NLP technologies for the underrepresented Khowar language while supporting language documentation, preservation, and future multilingual and multimodal research.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($250.00) platform fee paid to MDC.
Licensing
MDC Data License Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, facial recognition, voice-print matching, or any other technique intended to link recordings to real individuals.
Forbidden Usage
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, facial recognition, voice-print matching, or any other technique intended to link recordings to real individuals.
Khowar (ISO 639-3 code khw) is a Dardic language spoken primarily in Chitral district in Khyber Pakhtunkhwa, Pakistan, with additional speaker communities in Gilgit-Baltistan and among diaspora in cities such as Peshawar and Karachi. It has a well-developed literary tradition relative to many other Dardic languages, including active poetry, folklore, and language-preservation organizations, yet it remains under-resourced for speech and NLP technologies. This dataset includes a notable number of clips from speakers directly engaged in Khowar language and literature preservation work.
The dataset comprises unscripted, first-person interviews and talks in Khowar featuring speakers from Chitral and surrounding communities. Topics span personal and educational journeys (life in primary school, university life, higher education, M.Phil research, admission in law college, hostel life), professional life (career shifts such as moving to Peshawar for a new job, office work experiences, work for Khowar language and literature preservation), cultural heritage (the Pathak Festival, Khowar folktales including a recurring Fairy of the Sky and Eagle tale series, wedding customs of Chitral, traditional dishes such as tarbrad, village life in Bumborate and Kosht), social and environmental topics (proper use of water, mitigation of environmental disasters, heat waves), and closing reflections and community messages (favourite hobbies and sports, cherished childhood memories, the Chitral public library, and awards for language development).
The dataset consists of 152 folders, each corresponding to one recorded video segment. Each folder contains:
1 MP4 file — the video/audio recording of the segment
1 WAV file — the audio recording of the segment, extracted from the video
1 ELAN (.eaf) file — time-aligned annotation containing, among other fields:
Native Khowar transcription (tx_khw_native)
English translation (tr_eng)
In addition to the 152 folders, a separate .tsv metadata file accompanies the dataset, providing speaker-level and clip-level metadata for every recording, including:
Recording ID / folder name
Duration
Topic/description of the clip's content
Number of speakers in the clip
Speaker gender
Speaker age group
| Serial No. | Video Name | Duration | Topic / Description | Speakers Count | Speakers Gender | Speakers Age Group |
|---|---|---|---|---|---|---|
| 1 | 20260219-khw001 | 0:03:40 | Yesterday's activity | 2 | Male | 45 and Above |
| 2 | 20260219-khw002 | 0:02:33 | Yesterday's activity | 2 | Male | 45 and Above |
| 3 | 20260219-khw003 | 0:02:49 | Disadvantage of social media | 1 | Male | 45 and Above |
| 4 | 20260220-khw004 | 0:01:28 | Local games of my childhood | 1 | Male | 36-45 |
| 5 | 20260220-khw005 | 0:01:11 | My favourite sports | 1 | Male | 36-45 |
| 6 | 20260223-khw006 | 0:01:40 | Tradition of my village | 1 | Male | 45 and Above |