Release Date: 8/10/2026
Format: PARQUET
Size: 105.73 KB
Share
The Pakistan HSSC-I Mathematics Dataset is an instruction-following dataset derived from Pakistan Higher Secondary School Certificate Part-I (HSSC-I / Grade 11) mathematics content. The dataset consists of mathematics questions converted into instruction-response examples suitable for supervised fine-tuning (SFT) and educational natural language processing research. Each record contains an instruction, a mathematics problem, and a corresponding response or solution. Some prompts include phrases such as "think step by step" as part of the text format. Users should not assume these represent verified human reasoning unless independently validated. The dataset was prepared by curating Grade 11 mathematics problems, converting them into a consistent instruction-following format, performing basic cleaning and normalization, and exporting the final dataset as Apache Parquet files with predefined train, validation, and test splits. This dataset is intended for research and educational purposes. While reasonable care was taken during preprocessing, errors or inconsistencies may remain. Users should independently verify mathematical correctness before using the dataset in educational or assessment settings.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license. Commercial use is not permitted without obtaining appropriate permission. Any derivative datasets or redistributed versions must provide attribution and be shared under the same license. Users are responsible for ensuring compliance with applicable copyright laws and third party intellectual property rights.
Forbidden Usage
Its usage is forbidden too commercial use that is not permitted under the CC BY-NC-SA 4.0 license. Redistribution without proper attribution. Removal or alteration of licensing or attribution information. Use of the dataset in violation of applicable copyright or intellectual property laws. Presenting generated answers as authoritative educational material without independent verification.
The Pakistan HSSC-I Mathematics Dataset is an instruction-following dataset built from Higher Secondary School Certificate Part-I (HSSC-I / Grade 11) mathematics problems. The dataset is intended to support research in educational AI, supervised fine-tuning (SFT), mathematical reasoning, question answering, retrieval-augmented generation (RAG), and natural language processing.
11th-maths/
└── data/
├── train-00000-of-00001.parquet
├── validation-00000-of-00001.parquet
└── test-00000-of-00001.parquet
Each record is stored as a text based instruction following example suitable for supervised fine-tuning.
Typical records include:
An instruction
A mathematics problem
A response containing a solution or answer
Example:
### Instruction:
Help me answer this question.
Find the value of sin (1005°).
### Response:
Using the periodicity of the sine function, reduce the angle by multiples of 360° and evaluate the resulting expression.
Another example:
### Instruction:
Find the sum of the first 20 terms of the geometric progression
0.15, 0.015, 0.0015 ...
### Response:
Apply the finite geometric series formula using the given first term, common ratio, and number of terms.
The dataset was created by curating Grade 11 mathematics questions from educational materials available to the dataset creator and transforming them into a consistent instruction-following format suitable for machine learning research.
The preparation process included:
collecting mathematics questions;
converting them into instruction-response examples;
organizing examples into a consistent schema;
performing basic formatting and normalization;
exporting the data to Apache Parquet format;
creating predefined train, validation, and test splits.
Compensated · similar languages
A large multimodal dataset of actions performed in Nigeria: 10 egocentric videos and 1,666 time-stamped action-object-tool segments in the "health" domain.
A Hazargi-English parallel corpus derived from transcribed and translated video recordings