Release Date: 7/24/2026
Format: Parquet
Size: 20.26 MB
Share
This dataset is an Urdu language translation of the Stanford Alpaca instruction tuning dataset, filtered to 28,910 rows. Each entry contains an instruction, optional input, and output translated into Urdu, covering a broad range of task types including question answering, summarization, creative writing, classification, and reasoning. The dataset is designed to support instruction-tuning of Urdu language large language models, extending Alpaca style instruction data to a low resource language.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is a translation derivative of the Stanford Alpaca dataset, which itself carries a non commercial restriction due to its origin in OpenAI generated completions. Accordingly, this dataset is licensed under CC-BY-NC-SA-4.0: non-commercial use only, attribution required, and any redistributed derivative must be released under the same license.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission. Redistribution under a more permissive license than CC-BY-NC-SA-4.0. Use of this dataset to build models intended to directly compete with commercial LLM API providers, consistent with the restrictions on the original Alpaca source data.
Total samples: 28,910 (filtered)
Fields: urdu_instruction, urdu_input, urdu_output
Total size: 23.8 MB
This dataset is an Urdu translation of the Stanford Alpaca instruction-tuning dataset. The original Alpaca dataset was generated using OpenAI's API and is licensed CC-BY-NC-4.0 due to OpenAI's usage terms restricting commercial use of API-generated data.
Released under CC-BY-NC-SA-4.0 to reflect the inherited non-commercial restriction from the source Alpaca data, with an added ShareAlike requirement for derivatives.