Release Date: 8/10/2026
Format: PARQUET
Size: 1.92 MB
Share
The Urdu Sindhi Grapheme-to-Phoneme Dataset is a bilingual linguistic resource containing written Urdu and Sindhi words paired with their phonetic pronunciations in IPA notation. The dataset combines 30,657 Urdu grapheme-to-phoneme pairs and 99,676 Sindhi grapheme-to-phoneme pairs, for a total of 130,333 records. The Urdu portion contains train, validation, and test splits, while the Sindhi portion contains a train split. The dataset is intended to support research and development in grapheme-to-phoneme conversion, text-to-speech, automatic speech recognition, pronunciation modeling, speech synthesis, and computational phonology for Urdu and Sindhi.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is provided for non-commercial research, educational, academic, and other non-commercial purposes only. Users must comply with the CC-BY-NC-SA-4.0 license and provide appropriate attribution to the dataset and its sources. Commercial use, sale, licensing for commercial purposes, or incorporation into commercial products or services is not permitted without separate permission from the applicable rights holders.
Forbidden Usage
The dataset must not be used for commercial purposes, including commercial model training, commercial TTS or ASR services, paid APIs, or commercial products, unless separate permission has been obtained from the applicable rights holders. Users must not misrepresent the source or provenance of the data, remove attribution information, or redistribute modified versions under terms incompatible with CC-BY-NC-SA-4.0. The dataset must not be used for unlawful activities or in ways that violate applicable intellectual-property, copyright, privacy, or other legal requirements.
The Urdu-Sindhi Grapheme-to-Phoneme Dataset is a bilingual linguistic resource containing Urdu and Sindhi words paired with their phonetic pronunciations in IPA notation.
The dataset was created by collecting and processing language entries with available IPA pronunciations from Wiktionary, with the goal of providing a structured resource for low-resource Urdu and Sindhi speech and language technologies.
The combined dataset contains 130,333 grapheme-to-phoneme records:
Urdu: 30,657 records
Sindhi: 99,676 records
The dataset is distributed in Parquet format.
urdu_sindhi_g2p/
├── urdu-g2p/
│ ├── train.parquet
│ ├── validation.parquet
│ └── test.parquet
└── SND_G2P/
└── train.parquet
The Urdu portion contains:
graphemes — Urdu written form
phonemes — corresponding phonetic representation
Splits:
Train
Validation
Test
The Sindhi portion contains:
Grapheme — Sindhi written form
Phoneme — corresponding IPA pronunciation
Split:
Train
| Grapheme | Phoneme |
|---|---|
| ٹیٹ | ʈˈeːʈ |
| اوجی | ˈoːɟi |
| العین | alaˈiːn |
| گول | ɡˈoːl |
| معمول | mˈaːmuːl |
| دیانتدار | djˌaːnətdˈaːr |
| انٹیلی | ɪnʈˈeːli |
| دینا | dˈeːnaː |
| بجانا | bəɟˈaːnaː |
| مشاہیر | mˌəʃaːhˈiːr |
| ملوث | mʊlˈaʋʋas |
| دعا | dˈʊaː |
| کھلی | kʰˈʌli |
| مانسون | mˈaːnsoːn |
| برازیل | bˌəraːzˈiːl |
| ساتوں | sˈaːtõː |
| تیغ | tˈeːɣ |
| زمیں | zˈamĩː |
| بصیرت | bəʂˈiːrət |
| جنیفر | ɟənˈiːfər |
| Grapheme | Phoneme |
|---|---|
| ء | ɦ ə m z aː |
| آئون | ɑː ũː |
| آئينو | aː iː n oː |
| آبشار | aː b ʂ aː ɾ ə |
| آسمان | aː s m aː n ʊ |
| آفيس | aː f iː s |
| آنو | aː n oː |
| آنڊو | aː ɳ ɖ oː |
| آچر | aː t͡ʃ ə ɾ ʊ |
| آکاڙو | aː kʰ aː ɽ oː |
| اجازت | ɪ d͡ʒ aː z ə t̪ ə |
| اداس | ʊ d̪ ɑː s |
| اربع | ə ɾ b a (ː ) |
| اسان | ə s ãː |
| اسلام آباد | ɪ s l ɑː m ɑː b ɑː d̪ |
| اسپتال | ə s p ə t̪ ɑː l |
| الله حافظ | ə l ɦ aː f ɪ z |
| امتحانن | ɪ m t̪ ɪ ɦ aː n ə n |
| اميد | ʊ m eː d ə |
| انار | ə n aː r ʊ |
| اناناس | ə n aː n aː s |
| جانور | d͡ʒ ɑː n ʋ ə ɾ |
| جبل | d͡ʒ ə b ə l ʊ |
| جنوري | d͡ʒ ə n w ə ɾ iː |
A grapheme represents the written form of a word, while a phoneme represents its pronunciation.
For example:
Urdu
Grapheme: دعا
Phoneme: dˈʊaː
and:
Sindhi
Grapheme: آسمان
Phoneme: aː s m aː n ʊ
These mappings can be used as supervised examples for training G2P models that learn to convert written language into phonetic representations.
The dataset was collected from Wiktionary entries containing IPA pronunciation information.
For Urdu, entries were collected from Wiktionary's category for Urdu terms with IPA pronunciation. This category contains Urdu terms for which pronunciation is provided in IPA form.
For Sindhi, entries were collected from Wiktionary's category for Sindhi terms with IPA pronunciation. The collection process identifies Sindhi entries containing pronunciation information and extracts the written term together with its corresponding IPA representation.
The collected entries were then transformed into structured grapheme-to-phoneme pairs.
The processing pipeline consisted of:
Identifying Urdu and Sindhi Wiktionary entries containing IPA pronunciation.
Extracting the written word or grapheme.
Extracting the corresponding IPA pronunciation.
Removing unrelated dictionary information.
Converting the extracted records into structured tabular data.
Separating the data by language.
Creating train, validation, and test splits for the Urdu component.
Storing the resulting datasets in Parquet format.
Packaging the language-specific Parquet files into a single .tar.gz archive.
Only the grapheme-to-phoneme data required for the dataset is included in the released archive.
The pronunciation fields use IPA-style phonetic notation. Depending on the source entry, representations may contain:
IPA consonants
IPA vowels
Long-vowel markers
Nasalization
Stress marks
Affricates
Diacritics
Other phonological symbols
The pronunciation representation follows the information available in the corresponding Wiktionary entries.
| Language | Records | Splits |
|---|---|---|
| Urdu | 30,657 | Train / Validation / Test |
| Sindhi | 99,676 | Train |
| Total | 130,333 | — |
The dataset is intended to support non-commercial and academic research in:
Grapheme-to-Phoneme (G2P) conversion
Text-to-Speech (TTS)
Automatic Speech Recognition (ASR)
Pronunciation modeling
Speech synthesis
Multilingual speech technology
Computational phonology
Urdu NLP
Sindhi NLP
Low-resource language research
The dataset is derived from community-maintained Wiktionary entries. Consequently, pronunciation coverage and accuracy may vary between words and languages.
Users should validate the phonetic representations for their intended downstream application before using the dataset in production systems.
The dataset should be considered a structured research resource rather than a guaranteed authoritative pronunciation dictionary.
The primary source for the linguistic information is Wiktionary.
Wiktionary category:
Category:Urdu terms with IPA pronunciation
Wiktionary category:
Category:Sindhi terms with IPA pronunciation
The source entries are community-maintained and may be updated over time.
The underlying Wiktionary textual content is available under the Creative Commons Attribution-ShareAlike (CC BY-SA) licensing terms.
Recommended license: CC-BY-SA-4.0
Users should review the applicable Wiktionary licensing terms and attribution requirements before redistributing or publishing derivatives of this dataset.
This dataset was constructed from linguistic information available on Wiktionary.
Please attribute the underlying source to:
Wiktionary contributors
Source:
Urdu source category:
https://en.wiktionary.org/wiki/Category:Urdu_terms_with_IPA_pronunciation
Sindhi source category:
https://en.wiktionary.org/wiki/Category:Sindhi_terms_with_IPA_pronunciation
The complete dataset is distributed as:
urdu_sindhi_g2p.tar.gz
The archive contains only the dataset files:
urdu_sindhi_g2p/
├── urdu-g2p/
│ ├── train.parquet
│ ├── validation.parquet
│ └── test.parquet
└── SND_G2P/
└── train.parquet
No source-control metadata, cache files, hidden files, or unrelated auxiliary files are included in the release archive.
The dataset is provided for research and educational purposes. Users are responsible for verifying pronunciation data and ensuring that their use, redistribution, and derivative works comply with the applicable source licenses and terms.
The maintainers of this dataset do not claim ownership of the underlying Wiktionary contributions.