Release Date: 8/5/2026
Format: CSV
Size: 5.09 MB
Share
The Sindhi Transliteration Dataset by Proxima AI is a collection of Sindhi language sentences paired with their corresponding Roman Sindhi transliterations. The dataset was created and annotated by Proxima AI's team to support research and development in Sindhi natural language processing, particularly automatic transliteration between Sindhi Perso-Arabic script and Roman Sindhi. The dataset contains manually reviewed text pairs covering everyday topics, common vocabulary, names, places, and simple sentence structures. It is intended for training, fine-tuning, benchmarking, and evaluating transliteration models. Since Roman Sindhi does not have a universally accepted spelling standard, some transliterations reflect pronunciation or internal annotation guidelines and may vary from other Romanization conventions.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset may be used for non-commercial research, education, and machine learning in accordance with the CC BY-NC-SA 4.0 license. Users must provide appropriate attribution, indicate any modifications, and distribute derivative datasets under the same license.
Forbidden Usage
Commercial use without permission, redistribution without attribution, relicensing under an incompatible license, or any use that violates applicable laws or falsely implies endorsement by Proxima AI is prohibited.
Ethical Review
The dataset was internally reviewed by Proxima AI to ensure annotation quality and verify that it does not intentionally contain personal, confidential, or sensitive information.
Intended Use
This dataset is intended for non-commercial research, education, benchmarking, and the development of Sindhi transliteration systems, multilingual NLP applications, and language technology.
The Sindhi Transliteration Dataset by Proxima AI is a collection of aligned sentence pairs containing Sindhi text written in the Perso-Arabic script and its corresponding Roman Sindhi transliteration.
The dataset was created and reviewed by Proxima AI's team to support research and development in Sindhi natural language processing, particularly automatic transliteration between Sindhi Perso-Arabic script and Roman Sindhi.
The dataset contains 100,001 aligned sentence pairs, equivalent to approximately 100,000 records. The original Sindhi sentences were created using a hybrid workflow combining manually authored content with synthetically generated content produced in-house.
The synthetic and manually written sentences were designed to cover everyday topics, common vocabulary, personal names, locations, objects, actions, and simple sentence structures. Each Sindhi sentence was paired with a Roman Sindhi transliteration and reviewed by Proxima AI's annotation team to improve quality and consistency.
Because Roman Sindhi does not have a universally accepted spelling standard, some transliterations follow Proxima AI's internal annotation guidelines and may differ from other Romanization conventions.
Total sentence pairs: 100,001
Approximate number of records: 100,000
Languages: Sindhi and Roman Sindhi
Sindhi language code: snd
File format: CSV
Character encoding: UTF-8
Number of data columns: 2
The dataset size refers to the number of aligned sentence-pair records. It should not be interpreted as a token count. A separate token count has not been calculated because token totals depend on the tokenizer used.
The dataset is distributed as a UTF-8 encoded CSV file.
| Field | Type | Description |
|---|---|---|
sindhi_text | String | A sentence written in Sindhi Perso-Arabic script |
roman_sindhi_text | String | A Roman Sindhi transliteration of the same sentence |
The preferred column order is:
sindhi_text,roman_sindhi_text
Each row represents one aligned Sindhi and Roman Sindhi sentence pair.
The following examples illustrate the dataset schema and transliteration style.
| sindhi_text | roman_sindhi_text |
|---|---|
| سونيا تازي سبزي کائي ٿي | Sonia tazi sabzi khai thi |
| نويد پراڻي گهر ڏانهن وڃي ٿو | Naveed purane ghar dahn wanje tho |
| بشريٰ ننڍو ناول پڙهي ٿي | Bushra nandho novel parhe thi |
| حمزه تازو گوشت کائي ٿو | Hamza tazo gosht khai tho |
| حرا خوبصورت مسجد ڏانهن وڃي ٿي | Hira khoobsurat masjid dahn wanje thi |
| فاروق نئون سوڊا پيئي ٿو | Farooq naon soda pie tho |
| حسن صاف ڳوٺ ڏانهن وڃي ٿو | Hasan saaf goth dahn wanje tho |
| حسين مشھور سبق پڙهي ٿو | Hussain mashhoor sabaq parhe tho |
| ندا نئون صوف کائي ٿي | Nida naon soof khai thi |
| سليم عجيب ڳوٺ ڏانهن وڃي ٿو | Saleem ajeeb goth dahn wanje tho |
The original Sindhi sentences were produced using a hybrid data-creation workflow.
Some sentences were manually authored by Proxima AI's team, while others were synthetically generated in-house. The synthetic generation process was used to increase vocabulary coverage, sentence diversity, and the representation of common grammatical and sentence patterns.
The generated and manually authored content includes examples involving:
Everyday activities
Common objects and vocabulary
Personal names
Places and locations
Food and household items
Descriptive words
Frequently used verbs
Simple sentence structures
Each Sindhi sentence was paired with a corresponding Roman Sindhi transliteration. The resulting sentence pairs were reviewed by Proxima AI's annotation team to identify formatting issues, obvious language errors, and inconsistencies in transliteration.
The Roman Sindhi transliterations were prepared and reviewed using Proxima AI's internal annotation guidelines.
The review process focused on:
Ensuring that each Roman Sindhi entry corresponded to the associated Sindhi sentence
Checking the readability of Roman Sindhi forms
Improving spelling consistency across similar words
Identifying malformed or incomplete records
Confirming the expected CSV structure
Removing content that was clearly unsuitable for the dataset
The dataset is described as manually reviewed rather than fully manually authored because part of the original Sindhi content was created synthetically.