Release Date: 7/21/2026
Format: WAV, TSV
Size: 6.02 GB
Share
This dataset is a high-quality Single Speaker Urdu Text-to-Speech (TTS) corpus designed to support the development and evaluation of Urdu speech synthesis systems. It contains professionally recorded speech from a single native Urdu speaker, paired with accurate text transcriptions to provide consistent speech-text alignment. The recordings cover a diverse range of words, phrases, and sentences, helping models learn natural Urdu pronunciation, rhythm, and intonation while maintaining a consistent speaking style. The dataset is suitable for training neural TTS models, voice generation systems, speech technology research, and benchmarking speech synthesis performance in Urdu. By providing a reliable, single-speaker corpus for a low-resource language, this dataset contributes to the advancement of accessible and high-quality Urdu speech technologies for research and educational purposes.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($125.00) platform fee paid to MDC.
Licensing
MDC Data Licence Agreement 1.0
https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/Restrictions/Special Constraints
The downloader is restricted from using this dataset in any way that contravenes the MDC Data Licence Agreement 1.0
Forbidden Usage
The downloader is restricted from using this dataset in any way that contravenes the MDC Data Licence Agreement 1.0 You agree not to attempt to determine the identity of any speaker in the dataset. It is forbidden to use this dataset for voice cloning, biometric identification, surveillance, or the creation of synthetic voices mimicking speakers.
Intended Use
Training speech synthesis systems
Urdu is an Indo-Aryan language and the national language of Pakistan, spoken by millions of people worldwide. It is widely used in education, media, literature, and everyday communication. Urdu has a rich linguistic heritage and plays an important role in the development of speech and language technologies for low-resource languages.
Urdu (Perso-Arabic script): ا، ب، پ، ت، ٹ، ث، ج، چ، ح، خ، د، ڈ، ذ، ر، ڑ، ز، ژ، س، ش، ص، ض، ط، ظ، ع، غ، ف، ق، ک، گ، ل، م، ن، ں، و، ہ، ھ، ء، ی، ے۔
Speaker-1 (ID: urd1): Single native Urdu speaker (Female).
File type: WAV (Studio Quality)
Bit-depth: 16-Bit PCM
Audio sampling rate: 44.1 kHz
Background noise: ≤ −80 dBFS (typical average)
Reverberation time (RT60): Estimated ~0.18–0.25 seconds
Channel: Mono
Signal-to-noise ratio (SNR): ~65.9 dB (average)
| Domain | Clip count |
|---|---|
| Literature | 4512 |
| Travel-&-Tourism | 2017 |
| Informal-Urdu | 561 |
| Educational-Content | 553 |
| Cultural-Content | 2237 |
| Social-Media-Style | 366 |
| Health-&-Awareness | 402 |
| Religious | 1675 |
| News-&-Current-Affairs | 814 |
| General-Knowledge | 157 |
| Technology-&-Digital-Literacy | 941 |
| Story-Telling | 1295 |
| Drama-Scripts | 400 |