Release Date: 8/27/2026
Format: MP4, WAV, EAF
Size: 251.41 GB
Share
The Pashto Multimodal Dataset is a comprehensive language resource designed to support research in speech and language technologies for the Pashto language. It contains synchronized video and audio recordings accompanied by time-aligned transcriptions, speech segmentation, and detailed ELAN annotation files, enabling fine-grained multimodal analysis. The dataset is suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), speech segmentation, forced alignment, speaker and gesture analysis, multimodal language understanding, and linguistic research. By combining audiovisual data with rich annotations, it provides valuable resources for studying the interaction between speech, facial expressions, and other non-verbal communication cues. This dataset aims to advance the development of AI and NLP technologies for the underrepresented Pashto language while supporting language documentation, preservation, and future multilingual and multimodal research.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($250.00) platform fee paid to MDC.
Licensing
MDC Data License Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, facial recognition, voice-print matching, or any other technique intended to link recordings to real individuals.
Forbidden Usage
Any attempt to identify, re-identify, or expose the identity of speakers featured in this dataset is strictly forbidden. This includes cross-referencing with other datasets, facial recognition, voice-print matching, or any other technique intended to link recordings to real individuals.
Pashto (this collection uses Northern Pashto, ISO 639-3 code pbu) is spoken by tens of millions of people, primarily in Khyber Pakhtunkhwa (KP) province of Pakistan and across Afghanistan, with speakers in this dataset drawn from areas such as Swabi, Swat, and Charsadda. Pashto has a strong oral and literary tradition, including a well-known body of proverbs, poetry, and the Pashtunwali code of customs and social conduct. While it has a comparatively larger body of digital and broadcast content than some regional languages, spoken, dialect-rich, informal Pashto — as opposed to standard/written registers — remains under-resourced for speech and NLP research.
The dataset comprises unscripted, first-person interviews and talks in Pashto featuring speakers identified by speaker IDs, drawn from various towns and villages in Khyber Pakhtunkhwa. Topics span daily routines and childhood memories (childhood games, local games, memories of school and hostel life), place and identity (birthplace descriptions such as Swabi and Charsadda, village life in Manglor Swat), tradition and culture (Pashtun traditions and customs, wedding and death rituals, the role and impact of the Hujra, Pashto proverbs and poetry, friendship and rivalries among Pashtuns), social and generational topics (past vs. present education for women, the role of mother tongue in education, the impact of technology and artificial intelligence, healthy habits, children's care and training), religion (preparing for Namaz, the positive impact of prayer), and everyday life (professions, hobbies, travel experiences, weather's impact on daily life, environmental pollution).
The dataset consists of 209 folders, each corresponding to one recorded video segment. Each folder contains:
1 MP4 file — the video/audio recording of the segment
1 ELAN (.eaf) file — time-aligned annotation containing:
Native Pashto transcription (tx_pbu_native)
English translation (tr_eng)
In addition to the 209 folders, a separate TSV metadata file accompanies the dataset, providing speaker-level and clip-level metadata for every recording, including:
Recording ID / folder name
Speaker ID
Duration
Topic/description of the clip's content
Number of speakers in the clip
Speaker gender
Speaker age group
| Serrial No. | Video ID | Speaker ID | Duration | Topic / Description | # of Speakers | Gender | Age Group |
|---|---|---|---|---|---|---|---|
| 1 | 20260330-pbu001 | MPS101 | 0:01:00 | Talk about childhood games | 1 | Male | 18-25 |
| 2 | 20260330-pbu002 | MPS102 | 0:03:21 | Place of Birth (Swabi) | 1 | Male | 36-45 |
| 3 | 20260330-pbu003 | MPS102 | 0:01:43 | About childhood | 1 | Male | 36-45 |
| 4 | 20260330-pbu004 | MPS102 | 0:02:13 | Talk about festival (Eid ul Fitr) | 1 | Male | 36-45 |
| 5 | 20260331-pbu005 | MPS103 | 0:02:39 | Talk about what you did yesterday from morning to night | 1 | Male | 26-35 |
| 6 | 20260331-pbu006 | MPS104 | 0:01:14 | Discuss plans for the coming week | 1 | Male | 26-35 |