Release Date: 10/8/2026
Format: WEBM, TSV, JSON, JSONL, MD, TXT
Size: 1.33 GB
The Pashto Audio-Visual Speech Corpus, developed by Pashto AI and the Digital Pashto Association, is a multimodal speech dataset designed to support the development of speech and language technologies for Pashto. The corpus contains 9,421 validated sentence-level recordings, totaling 12.67 hours of audio, with each record pairing spoken Pashto audio, mouth-region video, and the corresponding Pashto transcription. Data was collected using MoLiÈRe from exactly nine male and female contributors, aged 20–35, who are educated Pashto speakers from different urban areas of Afghanistan. The corpus is intended for audio-visual speech recognition, lip reading, automatic speech recognition, multimodal Pashto AI, pronunciation and phonetic research, and experiments on speech recognition under noisy conditions. The original export contained 11,885 mapped record occurrences. After duplicate consolidation, 9,499 unique audio-video-text pairs remained. Of these, 78 records with invalid media were excluded, resulting in the final 9,421-record corpus. An additional 50 transcript-only records are retained as supplementary metadata. Media decoding was used to verify technical file readability, but does not guarantee transcription accuracy or audio-video lip synchronization. The dataset is released under the CC0 1.0 Universal Public Domain Dedication, making it openly available for anyone to use, copy, modify, redistribute, analyze, and build upon for research, commercial, educational, or other purposes without seeking permission, to the extent permitted by law. Our goal is to make Pashto language data as openly accessible as possible and support wider research, innovation, and the development of Pashto language technologies. Because the corpus contains only nine educated urban speakers aged 20–35, it should not be considered representative of all Pashto dialects, regions, age groups, or demographic populations.
Restrictions/Special Constraints
No restriction from our side
Forbidden Usage
No forbidden from our side
Intended Use
Lip reading; audio-visual and audio-only automatic speech recognition; multimodal Pashto AI; experiments on speech recognition under acoustic noise; and research on pronunciation, phonetics and speaker variation. Research performance and demographic representativeness have not been established.
Source folder: https://drive.google.com/drive/folders/1i-3cPHbHSx7IemuK5T5NXvKg805U9r-F Collection tool: MoLiÈRe, as reported by the creator.
4,421 retained pairs differ in audio/video duration by more than 0.5 seconds. This flags a need for alignment review and does not establish the cause or amount of any lip-sync offset. Full decoding checks file readability only; independent listening, transcript accuracy and lip-sync review remain outstanding.
Approximately 10 speakers and male/female participation are creator-reported, not independently verified. Seven source folders supplied retained recordings; folders are provenance groups, not verified speaker identities. No demographic inference was performed.
The public redistribution license and consent documentation await owner review. Prepared as a private draft; no review submission or publication has been performed.