Release Date: 7/21/2026
Format: TXT
Size: 592.14 KB
Share
This corpus contains a collection of traditional Kazakh proverbs (maqal-matelder) digitized from volumes 65-69 of the definitive "Babalar sözi" academic compilation published by Foliant in Astana. The dataset provides clean data designed to catalyze NLP research for the Kazakh language. Because proverbs are rich in cultural realias, historical context, and idiomatic phrasing, this corpus is uniquely positioned to serve as a foundation for "cultural competence" benchmarks, allowing researchers to evaluate an LLM's understanding of language-specific cultural nuances (e.g., through cloze tests or masked word evaluations). To ensure strict adherence to public domain copyright laws, all editorial apparatus has been excluded, and the original compiler's thematic ordering has been programmatically randomized.
Restrictions/Special Constraints
None
Forbidden Usage
None
This corpus contains a collection of traditional Kazakh proverbs (maqal-matelder). The texts were digitized from volumes 65-69 of the definitive "Babalar sözi" (Words of the Ancestors) academic series, published in Astana (2010-2011).
Beyond cultural preservation, this dataset is specifically designed for evaluating the "cultural competence" of Large Language Models (LLMs) in the Kazakh language. Proverbs contain dense cultural realias, idiomatic structures, and language-specific semantic dependencies. This corpus can be used to build fill-in-the-blank benchmarks (where specific words are omitted) to test whether an LLM actually understands regional cultural nuances rather than just generating statistically probable, yet culturally empty, translations.
Total Word Count: 171240
Files:
proverbs_vol65.txt: 35903 words (6153 proverbs)
proverbs_vol66.txt: 32822 words (5494 proverbs)
proverbs_vol67.txt: 39052 words (6883 proverbs)
proverbs_vol68.txt: 30016 words (5300 proverbs)
proverbs_vol69.txt: 33447 words (5833 proverbs)
The following academic volumes served as the source material for this corpus:
Alpysbaeva, K., Alibekov, T., & Kosan, S. (Comp.). (2010). Babalar sozi: Vol. 65. Qazaq maqal-matelderi. Astana: Foliant.
Alpysbaeva, K., Alibekov, T., & Elesbay, N. (Comp.). (2010). Babalar sozi: Vol. 66. Qazaq maqal-matelderi. Astana: Foliant.
Akimova, T., Oralbek, A., & Oryngali, K. (Comp.). (2011). Babalar sozi: Vol. 67. Qazaq maqal-matelderi. Astana: Foliant.
Alibekov, T., Akimova, T., Oralbek, A., & Oryngali, K. (Comp.). (2011). Babalar sozi: Vol. 68. Qazaq maqal-matelderi. Astana: Foliant.
Alpysbaeva, K., Alibekov, T., & Elesbay, N. (Comp.). (2011). Babalar sozi: Vol. 69. Qazaq maqal-matelderi. Astana: Foliant.
The dataset is provided in plain text across 5 files. Each file begins with a single YAML Front Matter block containing the compilation metadata. Following this header, the proverbs are listed sequentially, with each entry delimited by a blank line.
Below is a verbatim segment of the beginning of the file:
---
title: "Қазақ мақал-мәтелдері"
lang: "kk"
year: "Unknown"
source_original: "Unknown"
source_container: "Akimova, T., Oralbek, A., & Oryngali, K. (Comp.). (2011). Babalar sozi: Vol. 67. Qazaq maqal-matelderi [Words of the Ancestors: Vol. 67. Kazakh Proverbs]. Astana: Foliant."
---
ҚАЗАҚ МАҚАЛ-МӘТЕЛДЕРІ
Сыпайыны үйде көрме, түзде көр.
Бөспе атаққа қызығады, Бөрі тамаққа қызығады.
Таудай істің тарыдай түйіні бар.
Құтты қонақ қонса, Қой егіз табады. Құтсыз қонақ келсе, Қойға қасқыр шабады.
Келінің жақсы болса, Дәулет келді десеңші.
Note: To neutralize any copyright claims over the specific arrangement (thematic grouping or alphabetization) enacted by the compilers, the order of the proverbs within this digital corpus has been programmatically randomized.
This dataset was created through the following pipeline:
Acquisition: Sourced high-quality PDF scans of the Babalar sözi volumes from the adebiportal.kz literary portal.
Text Extraction: The single-column physical layouts of the volumes
were parsed using PyMuPDF to extract the raw text layer page by page.
LLM OCR & Verbatim Guardrails: Raw text layers were passed to
gemini-2.5-flash for extraction. To prevent hallucination or truncation,
a programmatic verify_verbatim guardrail utilizing look-ahead logic
enforced strict substring matching between the LLM output and the raw PDF
text layer before appending to the dataset.
Folklore Status: According to Article 8 of the Law of the Republic of Kazakhstan "On Copyright and Related Rights", works of folk art (folklore) are not subject to copyright. Therefore, the proverb texts themselves reside in the public domain. The editorial ordering of the original printed volumes has been scrambled to prevent derivative copyright claims on the arrangement.
Publisher Commentary: Scientific commentaries, introductions, and footnotes provided by the compilers/publishers are copyrighted material and have been strictly excluded from this corpus.
Usage: The texts provided are in the public domain (CC0-1.0), but we would appreciate attribution where you reference this Mozilla Data Collective dataset.