Release Date: 8/7/2026
Format: MP4, EAF
Size: 44.00 GB
Share
The Hazargi Multimodal Dataset is a comprehensive language resource designed to support research in speech and language technologies for the Hazargi language. It contains synchronized video and audio recordings accompanied by time-aligned transcriptions, speech segmentation, and detailed ELAN annotation files, enabling fine-grained multimodal analysis. The dataset is suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), speech segmentation, forced alignment, speaker and gesture analysis, multimodal language understanding, and linguistic research. By combining audiovisual data with rich annotations, it provides valuable resources for studying the interaction between speech, facial expressions, and other non-verbal communication cues. This dataset aims to advance the development of AI and NLP technologies for the underrepresented Hazargi language while supporting language documentation, preservation, and future multilingual and multimodal research.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($250.00) platform fee paid to MDC.
Licensing
MDC Data License Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, facial recognition, voice-print matching, or any other technique intended to link recordings to real individuals.
Forbidden Usage
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, facial recognition, voice-print matching, or any other technique intended to link recordings to real individuals.
Hazargi (also spelled Hazaragi) is a variety of Dari Persian spoken by the Hazara people, primarily in the Hazarajat region of central Afghanistan, as well as in diaspora communities in Pakistan (notably Quetta) and Iran. It is closely related to standard Dari but includes distinct vocabulary, including a number of Turkic and Mongolic loanwords reflecting the historical origins of the Hazara people. Hazargi is considered a low-resource language for natural language processing and speech technology, with very limited digital or annotated data currently available.
The dataset comprises unscripted, first-person interviews and talks in Hazargi featuring speakers from diverse backgrounds within the Hazara community including artists, musicians, athletes, doctors, students, entrepreneurs, activists, and tradespeople. Topics span personal introductions and life journeys, Hazara history and culture, language preservation, education, business and entrepreneurship, sports (notably boxing and karate), traditional crafts and cuisine, social issues (education access, gender barriers, drug use, media awareness), organizational profiles (e.g., Mechid, Deedar, Hazara Jirga, Panahi Musical Academy), and closing messages of advice or awareness directed at youth and the broader community.
The dataset consists of 181 folders, each corresponding to one recorded video segment. Each folder contains:
1 MP4 file — the video/audio recording of the segment
1 ELAN (.eaf) file — time-aligned annotation containing:
Native Hazargi transcription
English translation
In addition to the 181 folders, a separate Excel (.xlsx) metadata file accompanies the dataset, providing speaker-level and clip-level metadata for every recording, including:
Recording ID / folder name
Duration
Topic/description of the clip's content
Number of speakers in the clip
Speaker gender
Speaker age group
| Serial No. | Recording ID | Duration | Topic / Description | # of Speakers | Gender | Age Group |
|---|---|---|---|---|---|---|
| 1 | 20260215-haz001 | 0:04:51 | Self introduction and talks about his studies | 1 | Male | 45 and Above |
| 2 | 20260215-haz002 | 0:05:25 | Talks about Hazara history | 1 | Male | 45 and Above |
| 3 | 20260215-haz003 | 0:04:47 | Talks about history in general | 1 | Male | 45 and Above |
| 4 | 20260215-haz004 | 0:03:19 | Speaker's message to youngsters | 1 | Male | 45 and Above |
| 5 | 20260217-haz005 | 0:04:16 | How he started his work initially | 1 | Male | 45 and Above |
| 6 | 20260217-haz006 | 0:04:44 | About his work experiences and hardships | 1 | Male | 45 and Above |
(Full metadata for all 181 recordings is provided in the accompanying Excel and TSV files.)