Release Date: 8/5/2026
Format: PARQUET
Size: 436.48 MB
Share
UROCR is an Urdu Optical Character Recognition dataset compiled from multiple sources, containing 40,769 image-text pairs suitable for training and evaluating Urdu OCR models. The dataset spans a range of image types, scanned documents, printed text, synthetic text renders, and natural scene text and covers both Nastaliq and Naskh script styles. Text content spans multiple domains including literature, news, religious texts, and general text, with text length ranging from single words to full sentences. The dataset is split into train (32,615 / 80%), validation (4,077 / 10%), and test (4,077 / 10%) subsets.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is intended for non-commercial research use only, including training and evaluation of Urdu OCR models. Attribution to Proxima AI is required for any use.
Forbidden Usage
Use of this dataset, or of models trained on it, for commercial products or services is forbidden, consistent with the non-commercial (NC) restriction of this license. Redistribution without retaining attribution to Proxima AI. Use of scanned document images to attempt identification of any individual or institution referenced in the source documents.
Ethical Review
This dataset compiles image-text pairs from Proxima AI's own document scans, archival material, and team captured natural scene text, along with in house synthetic text renders and augmentations. No third party or externally sourced images were used, and no new human subject data collection was involved. Source material spans general literature, news, and religious domain text; no private, sensitive, or personally identifying documents are included.
Intended Use
This dataset is intended for non-commercial training and evaluation of Urdu OCR models across multiple text styles and image conditions (scanned, printed, synthetic, natural scene text).
Total samples: 40,769 image-text pairs
Train: 32,615 (80.0%)
Validation: 4,077 (10.0%)
Test: 4,077 (10.0%)
image: Image containing Urdu text (various resolutions)
text: Corresponding ground-truth Urdu text transcription
Scanned documents
Printed text images
Synthetic text images
Natural scene text
Scanned document and printed-text images were produced from Proxima AI's own document scans and archival material. Natural scene text images were captured or sourced by the Proxima AI team, with some samples produced via image augmentation of the original captures. Synthetic text renders were generated in-house. No third-party or externally scraped image sources were used, so no additional rights clearance was required for the source content.
Script: Urdu (Nastaliq and Naskh styles)
Domains: Literature, news, religious texts, general text
Text length: Ranges from single words to full sentences
Released under CC-BY-NC-SA-4.0.