Release Date: 7/22/2026
Format: PARQUET
Size: 357.97 KB
Share
Urdu Medical SFT is a safety-aligned dataset of 3,072 Urdu medical question-answer pairs designed for supervised fine-tuning (SFT) of large language models in the medical domain. The dataset spans 14 medical specialties, with General Medicine (51.9%), Psychiatry (13.4%), and Cardiology (8.7%) as the largest categories, split into 85% train (2,612 records) and 15% validation (460 records). Every response includes an embedded Urdu safety disclaimer, defers to licensed medical professionals, uses hedged (non-definitive) language, and avoids naming specific medications without a doctor qualifier. The dataset was constructed with explicit safety alignment as a primary design goal rather than raw medical knowledge coverage.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
Restrictions/Special Constraints
This dataset is intended for non-commercial research and fine tuning of medical domain language models only. Attribution to Proxima AI is required for any use. This dataset has not been clinically validated and must not be used as a substitute for professional medical review before deployment in any user facing application.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission from Proxima AI. Use of this dataset to deploy a medical chatbot or QA system directly to end users without human clinical review and additional safety validation. Redistribution without retaining attribution to Proxima AI . Removal or bypassing of the embedded safety disclaimers when reusing dataset content.
Total records: 3,072 (Train: 2,612 / 85%, Validation: 460 / 15%)
Average question length: 43.4 characters
Average answer length: 268.5 characters
Paraphrased variants: 76
| Level | Count | % |
|---|---|---|
| Basic | 501 | 16.3% |
| Intermediate | 2,075 | 67.5% |
| Advanced | 496 | 16.1% |
Covers 14 medical specialties, led by General Medicine (51.9%), Psychiatry (13.4%), and Cardiology (8.7%), with several specialties (Pulmonology, Orthopedics, Neurology, Ophthalmology) each under 2%.
Every response includes an embedded Urdu safety disclaimer (اہم نوٹ), references a licensed doctor or specialist, avoids definitive diagnoses through hedged language, and does not name specific medications without a doctor qualifier.
Domain imbalance toward General Medicine
Not clinically validated by licensed physicians
Modest scale (3,072 records) relative to English medical SFT datasets
Reflects general South Asian Urdu; may not capture all regional dialects
Released under CC-BY-NC-SA-4.0.