Release Date: 7/23/2026
Format: Parquet
Size: 18.94 MB
Share
This dataset is a Sindhi language translation of the Stanford Alpaca instruction tuning dataset, filtered to 28,910 rows. Each entry contains an instruction, optional input, and output triple translated into Sindhi, covering a broad range of task types including question answering, summarization, creative writing, classification, and reasoning. The dataset is intended to support instruction-tuning of Sindhi language large language models, extending Alpaca style instruction data to a low resource South Asian language.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is a translation derivative of the Stanford Alpaca dataset, which itself carries a non-commercial restriction due to its origin in OpenAI generated completions. Accordingly, this dataset is licensed under CC-BY-NC-SA-4.0: non-commercial use only, attribution required, and any redistributed derivative must be released under the same license.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission. Redistribution under a more permissive license than CC-BY-NC-SA-4.0. Use of this dataset to build models intended to directly compete with commercial LLM API providers, consistent with the restrictions on the original Alpaca source data.
Total samples: 28,910 (filtered from a larger translated set)
Fields: sindhi_instruction, sindhi_input, sindhi_output
Total size: 22.2 MB
This dataset is a Sindhi translation of the Stanford Alpaca instruction-tuning dataset, filtered for quality. The original Alpaca dataset was generated using OpenAI's text-davinci-003 API and is licensed CC-BY-NC-4.0 due to OpenAI's usage terms restricting commercial use of API generated data.
Released under CC-BY-NC-SA-4.0 to reflect the inherited non-commercial restriction from the source Alpaca data, with an added ShareAlike requirement for derivatives.