Release Date: 7/20/2026
Format: TXT
Size: 142.68 KB
Share
This corpus contains a collection of traditional Uzbek proverbs (maqollar) digitized from the definitive 2005 academic compilation published by Sharq in Tashkent. The dataset provides clean data designed to catalyze NLP research for the Uzbek language. Because proverbs are rich in cultural realias, historical context, and idiomatic phrasing, this corpus is uniquely positioned to serve as a foundation for "cultural competence" benchmarks, allowing researchers to evaluate an LLM's understanding of language-specific cultural nuances (e.g., through cloze tests or masked word evaluations). To ensure strict adherence to public domain copyright laws, all editorial apparatus has been excluded, and the original compiler's thematic ordering has been programmatically randomized.
Restrictions/Special Constraints
None
Forbidden Usage
None
This corpus contains a collection of traditional Uzbek proverbs (maqollar). The texts were digitized from the definitive 2005 academic volume published in Tashkent.
Beyond cultural preservation, this dataset is specifically designed for evaluating the "cultural competence" of Large Language Models (LLMs) in the Uzbek language. Proverbs contain dense cultural realias, idiomatic structures, and language-specific semantic dependencies. This corpus can be used to build fill-in-the-blank benchmarks (where specific words are omitted) to test whether an LLM actually understands regional cultural nuances rather than just generating statistically probable, yet culturally empty, translations.
Total Word Count: 51555
Files:
proverbs.txt: 51555 words (9150 proverbs)
The following academic volume served as the source material for this corpus:
Mirzayev, T., Musoqulov, A., & Sarimsoqov, B. (Comps.). (2005). O'zbek xalq maqollari [Uzbek Folk Proverbs]. Tashkent: Sharq.
The dataset is provided in plain text. The file begins with a single YAML Front Matter block containing the compilation metadata. Following this header, the proverbs are listed sequentially, with each entry delimited by a blank line.
Below is a verbatim segment of the beginning of the file:
---
title: "O'zbek xalq maqollari"
lang: "uz"
year: "Unknown"
source_original: "Unknown"
source_container: "Mirzayev, T., Musoqulov, A., & Sarimsoqov, B. (Comp.). (2005). O'zbek xalq maqollari [Uzbek Folk Proverbs]. Tashkent: Sharq."
---
O'ZBEK XALQ MAQOLLARI
Dono so'zini tergar, Nodon — ko'zini.
Gado to'ydi — qayqaydi.
Yomonning o'zi nima-yu, so'zi nima.
It qorni to'ygan uydan ketmas.
Urishqoq xotin bor qishloqqa qo'riqchining hojati yo'q.
Note: To neutralize any copyright claims over the specific "creative arrangement" (thematic grouping or alphabetization) enacted by the 2005 compilers, the order of the proverbs within this digital corpus has been programmatically randomized.
This dataset was created through the following pipeline:
Acquisition: Sourced high-quality PDF scans of the 2005 volume from the Ziyouz digital library.
Block-Level Extraction: Multi-column physical layouts were parsed
using PyMuPDF block-level coordinate sorting to ensure linear text
integrity.
LLM OCR & Verbatim Guardrails: Raw text layers were passed to
gemini-2.5-flash for extraction. To prevent hallucination, a
programmatic verify_verbatim guardrail enforced byte-for-byte
matching between the LLM output and the raw PDF text layer before
appending to the dataset.
Folklore Status: According to Article 8 of the Law of the Republic of Uzbekistan "On Copyright and Related Rights" (No. ZRU-42), works of folklore are not subject to copyright. Therefore, the proverb texts themselves reside in the public domain. The editorial ordering of the original printed volume has been scrambled to prevent derivative copyright claims on the arrangement.
Publisher Commentary: Scientific commentaries, introductions, and footnotes provided by the compilers/publishers are copyrighted material and have been strictly excluded from this corpus.
Usage: The texts provided are in the public domain (CC0-1.0), but we would appreciate attribution where you reference this Mozilla Data Collective dataset.