Release Date: 7/21/2026
Format: TSV
Size: 434.04 KB
Share
A parallel corpus of over 3,000 sentences (over 80k English tokens) of Hazargi and English. The sentences are derived from a multimodal corpus consisting of transcribed and translated Hazargi language videos.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($75.00) platform fee paid to MDC.
Licensing
MDC Data License Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
Any attempt to identify, re-identify, or expose the identity of speakers of the texts featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, or any other technique intended to link texts to real individuals.
Forbidden Usage
Any attempt to identify, re-identify, or expose the identity of speakers of the texts featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, or any other technique intended to link texts to real individuals.
Intended Use
Machine Translation
Hazargi (also spelled Hazaragi) is a variety of Dari Persian spoken by the Hazara people, primarily in the Hazarajat region of central Afghanistan, as well as in diaspora communities in Pakistan (notably Quetta) and Iran. It is closely related to standard Dari but includes distinct vocabulary, including a number of Turkic and Mongolic loanwords reflecting the historical origins of the Hazara people. Hazargi is considered a low-resource language for natural language processing and speech technology, with very limited digital or annotated data currently available.
The dataset is derived from a collection of unscripted, first-person interviews and talks in Hazargi featuring speakers from diverse backgrounds within the Hazara community including artists, musicians, athletes, doctors, students, entrepreneurs, activists, and tradespeople. Topics span personal introductions and life journeys, Hazara history and culture, language preservation, education, business and entrepreneurship, sports (notably boxing and karate), traditional crafts and cuisine, social issues (education access, gender barriers, drug use, media awareness), organizational profiles (e.g., Mechid, Deedar, Hazara Jirga, Panahi Musical Academy), and closing messages of advice or awareness directed at youth and the broader community.
The dataset contains a total of 3.3k English sentences, 86k English tokens, and 91k Hazargi tokens.
The dataset contains a single TSV file with three fields:
video_id: an identifier of the original video the clip was extracted from. This is also useful for lookup in the metadata file, discussed below.
en_sentence: an English translation text of the spoken Hazargi.
haz_sentence: the Hazargi transcriptions