Release Date: 7/24/2026
Format: PARQUET
Size: 1.34 GB
Share
The Sindhi OCR Dataset is designed for optical character recognition research and development in Sindhi (Arabic script). It contains 278,000 synthetically generated text-image pairs, where each image consists of black Sindhi text rendered on a white background using nine different Lateef and Nargis fonts. Text sources include single Sindhi dictionary words as well as two and three-word combinations to increase lexical diversity and improve model generalization. The dataset is split into train (256,000), validation (8,000), and test (14,000) subsets.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is licensed under CC-BY-NC-4.0 and is intended for non-commercial research use only, including training and evaluation of Sindhi OCR models. Attribution is required for any use. This dataset is synthetically generated and should not be used as a substitute for real-world scanned document evaluation without additional testing.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission. Redistribution without retaining attribution to Danish Mahdi. Representing model performance on this dataset as equivalent to real world scanned or handwritten Sindhi OCR performance without additional validation.
Ethical Review
This dataset consists entirely of synthetically generated text images using dictionary word combinations rendered in various fonts. No real scanned documents, handwriting samples, or personally identifying information were used in its creation, so no human-subject or privacy concerns apply.
Intended Use
This dataset is intended for non-commercial training and benchmarking of Sindhi OCR models, vision-language model fine-tuning, and low-resource OCR research. It is not recommended for direct deployment on real-world scanned or handwritten documents without additional validation.
Text source: Sindhi dictionary words, two-word combinations, three-word combinations
Image format: Black text on white background, synthetically generated
Ground truth: Unicode Sindhi text
Lateef-Regular, Lateef-Bold, Lateef-Light, Lateef-Medium, Lateef-SemiBold, Lateef-ExtraBold, Lateef-ExtraLight, MB-Lateefi-SKv2.0, MBNargisNew3.2
Train: 256,000 (92.1%)
Validation: 8,000 (2.9%)
Test: 14,000 (5.0%)
Total: 278,000
Sindhi OCR model training
Handwritten-to-printed OCR transfer research
Vision-language model fine-tuning
Text recognition benchmarking
Low-resource language OCR research
Real-world document OCR evaluation without additional testing
Historical document recognition
Handwritten Sindhi OCR
Synthetically generated; may not fully represent real-world scanned documents
Only black text on white backgrounds is included
Does not contain handwritten text
Document layouts, tables, stamps, signatures, and complex formatting are not represented
Released under CC-BY-NC-SA-4.0.