Release Date: 7/22/2026
Format: PARQUET
Size: 4.58 GB
Share
Syntheic Urdu Audios is a synthetically generated Urdu dataset of 7,963 audio-text pairs, pairing conversational transcript text with corresponding single speaker audio recordings. Content consists of casual, humorous multi turn dialogue exchanges covering everyday topics written in a natural, colloquial conversational Urdu register, including expressive stage directions (e.g. "laughing," "sighing") embedded in the text. The dataset is intended for training and evaluating Urdu TTS and ASR systems on natural conversational speech patterns rather than formal or scripted text.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is intended for non commercial research use only, including training and evaluation of Urdu TTS/ASR models on conversational speech. Attribution to Proxima AI is required for any use.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission from Proxima AI. You agree not to attempt to determine the real world identity of the speaker in this dataset. Any attempt to clone the speaker's voice for impersonation, fraud, or deceptive purposes is forbidden. Redistribution without retaining attribution to Proxima AI.
Total samples: 7,963 audio-text pairs
Fields: id, transcript, voice (single speaker: "amuch"), text, audio_path
Total size: 5.6 GB
Casual, colloquial Urdu conversational dialogue with expressive stage directions (e.g. laughing, sighing) embedded in transcripts. Topics include everyday humorous banter content is synthetically generated rather than sourced from real conversations.
Released under CC-BY-NC-SA-4.0.