Release Date: 7/23/2026
Format: Parquet
Size: 10.37 MB
Share
HH-RLHF Urdu is an Urdu-language translation of Anthropic's Helpful and Harmless RLHF (HH-RLHF) dataset, containing 11,225 preference pairs of "chosen" and "rejected" multi-turn conversations. Each example preserves the original turn structure and user/assistant roles, translated using TranslateGemma27B, a 27B-parameter instruction-tuned multilingual translation model. The dataset is intended to support reward model training and RLHF based fine-tuning of Urdu language LLMs, extending Anthropic's helpfulness/harmlessness preference framework to a low resource language.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is a translation of Anthropic's HH-RLHF dataset (originally MIT-licensed). The Urdu translation itself is released under CC-BY-NC-SA-4.0: non-commercial use only, attribution required, and any redistributed derivative must be released under the same license.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission. Redistribution under a more permissive license than CC-BY-NC-SA-4.0. Use of this dataset to train models intended to bypass or contradict the original helpfulness/harmlessness alignment goals.
Total rows: 11,225
Fields: chosen (preferred conversation), rejected (dispreferred conversation), chosen_score, rejected_score
Format: JSONL, one dict per line
Total size: 44.4 MB
Model: TranslateGemma27B, a 27B-parameter instruction-tuned Gemma variant specialized for high-quality multilingual translation
Process: Each conversation turn was translated independently while preserving the original turn structure and roles (user/assistant), prompted to produce fluent, natural Urdu without adding or omitting information
This dataset translates Anthropic's HH-RLHF dataset, introduced in "Constitutional AI: Harmlessness from AI Feedback" (Bai et al., 2022, arXiv:2212.08073). The original dataset is MIT-licensed.
Used in training mahwizzzz/Qalb-1.0-8B-DPO, a DPO-trained text generation model.
The Urdu translation is released under CC-BY-NC-SA-4.0. The original English HH-RLHF dataset remains MIT-licensed.