Release Date: 7/23/2026
Format: Parquet
Size: 14.80 GB
Share
Synthetic Urdu 2 is a synthetically generated Urdu dataset of 85,327 audio-text pairs, pairing poetic, philosophical, and spoken-word style verses with corresponding single-speaker audio recordings (voice type: "verse"). Text content is reflective and lyrical in nature, exploring themes of existence, spirituality, love, the universe, and self awareness in a literary register distinct from conversational speech. The dataset is intended for training and evaluating Urdu TTS/ASR models on expressive, poetic recitation style speech, complementing conversational datasets like synthetic urdu.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is licensed under CC-BY-NC-SA-4.0 and is intended for non-commercial research use only, including training and evaluation of Urdu TTS/ASR models on poetic/recitation-style speech. Attribution to Proxima AI is required for any use. Any derivative datasets or models built using this data that are themselves redistributed must be released under the same CC-BY-NC-SA-4.0 license (ShareAlike).
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission from Proxima AI. You agree not to attempt to determine the real-world identity of the speaker in this dataset. Any attempt to clone the speaker's voice for impersonation, fraud, or deceptive purposes is forbidden. Redistribution without retaining attribution to Proxima AI. Redistributing this dataset, or any derivative work built from it, under a different or more permissive license than CC-BY-NC-SA-4.0.
Total samples: 85,327 audio-text pairs
Fields: id, text, audio_path, voice (single voice type: "verse")
Total size: 20.5 GB
Poetic, reflective, spoken-word style Urdu text exploring themes of existence, spirituality, love, and the universe. Content is synthetically composed rather than sourced from existing published poetry.
This dataset complements [https://mozilladatacollective.com/datasets/cmrw6gnqc005pnv073f87ulq7] which covers conversational/humorous dialogue; Synthetic Urdu 2 focuses instead on poetic/recitation style speech to broaden stylistic coverage for Urdu TTS training.
Released under CC-BY-NC-SA-4.0. Non-commercial use only, attribution to Proxima AI required, and any redistributed derivative must be shared under the same license (ShareAlike).