Release Date: 7/22/2026
Format: TXT
Size: 86.48 KB
Share
This corpus contains a collection of traditional Kumyk proverbs and sayings (макъаллар ва айтывлар) extracted from the native digital PDF of the 2015 academic compilation by N. E. Gadzhiakhmedov published in Makhachkala. The dataset provides clean data designed to catalyze NLP research for the lesser-resourced Kumyk language. Because proverbs and sayings are rich in cultural realias, historical context, and idiomatic phrasing, this corpus is uniquely positioned to serve as a foundation for "cultural competence" benchmarks, allowing researchers to evaluate an LLM's understanding of language-specific cultural nuances (e.g., through cloze tests or masked word evaluations). To ensure strict adherence to public domain copyright laws, all Russian translations and editorial apparatus have been excluded via a typographical regex extraction pipeline, and the original compiler's alphabetical ordering has been programmatically randomized.
Restrictions/Special Constraints
None
Forbidden Usage
None
This corpus contains a collection of traditional Kumyk proverbs and sayings (макъаллар ва айтывлар). Kumyk is a Turkic language of the Kipchak branch, spoken primarily by the Kumyk people in the Republic of Dagestan (Russia). As a lesser-resourced language, digitized native texts are relatively scarce. The texts in this corpus were extracted from N. E. Gadzhiakhmedov's 2015 bilingual dictionary.
Beyond cultural preservation, this dataset is specifically designed for evaluating the "cultural competence" of Large Language Models (LLMs) in the Kumyk language. Proverbs and sayings contain dense cultural realias, idiomatic structures, and language-specific semantic dependencies. This corpus can be used to build fill-in-the-blank benchmarks (where specific words are omitted) to test whether an LLM actually understands regional cultural nuances rather than just generating statistically probable translations.
In this corpus, some entries contain words or phrases enclosed in parentheses. These represent valid, traditional variants of the expression. The text in the parentheses can be used to replace the word or phrase immediately preceding it.
Single word alternates:
Example: БИРЕВ КЪОЙ ИЗЛЕЙ, БИРЕВ ТОЙ ИЗЛЕЙ (ГЁЗЛЕЙ)
This means БИРЕВ КЪОЙ ИЗЛЕЙ, БИРЕВ ТОЙ ГЁЗЛЕЙ is also a valid expression.
Multiple comma-separated alternates:
Example: АЛМА ТЕРЕГИНДЕН АРИ (АРЕК, УЗАКЪГЪА) ТЮШМЕС
This means both АЛМА ТЕРЕГИНДЕН АРЕК ТЮШМЕС and АЛМА ТЕРЕГИНДЕН УЗАКЪГЪА ТЮШМЕС are valid expressions.
Phrase alternates:
Example: КЪОНАКЪ КЪОНАКЪНЫ СЮЙМЕС, УЬЙ ЕСИ ЭКЕВЮН ДЕ (БИРИН ДЕ) СЮЙМЕС
This means both ...УЬЙ ЕСИ ЭКЕВЮН ДЕ СЮЙМЕС and ...УЬЙ ЕСИ БИРИН ДЕ СЮЙМЕС are valid expressions.
Total Word Count: 26612
Files:
proverbs.txt: 26612 words (4931 entries)
The following academic volume served as the source material for this corpus:
Gadzhiakhmedov, N. E. (2015). Kumyksko-russkiy slovar' poslovits i pogovorok [Kumyk-Russian Dictionary of Proverbs and Sayings]. Makhachkala: DGU.
The dataset is provided in plain text. The file begins with a single YAML Front Matter block containing the compilation metadata. Following this header, the entries are listed sequentially, with each item delimited by a blank line.
Below is a verbatim segment of the beginning of the file:
---
title: "Къумукъ макъаллар ва айтывлар"
lang: "kum"
year: "Unknown"
source_original: "Unknown"
source_container: "Gadzhiakhmedov, N. E. (2015). Kumyksko-russkiy slovar' poslovits i pogovorok [Kumyk-Russian Dictionary of Proverbs and Sayings]. Makhachkala: DGU."
---
КЪУМУКЪ МАКЪАЛЛАР ВА АЙТЫВЛАР
НАМУССУЗ ГИШИ АКЪЧАГЪА АТАСЫН САТАР
КЁПНЮ АВЗУ КАРАМАТ
ИТ ГЬАПЛАР − КЕРИВАН ГЕЧЕР, КЁП СЁЙЛЕГЕН ШОРПА ИЧЕР
Note: To neutralize any copyright claims over the specific arrangement (thematic grouping or alphabetization) enacted by the 2015 compiler, the order of the proverbs and sayings within this digital corpus has been programmatically randomized.
This dataset was created through the following pipeline:
Acquisition: Sourced from a native digital PDF of the 2015 academic volume.
Algorithmic Extraction: The text was parsed using a purely algorithmic,
regex-based typographical script via PyMuPDF.
Filtering: The extraction logic leveraged the book's strict uppercase-proverb formatting to isolate the Kumyk folklore items while automatically bypassing all Russian translations, dictionary abbreviations, and editorial commentary.
Verification: The extracted corpus was manually proofread and inspected to ensure that all extraneous metadata, translation fragments, typographical anomalies (like footnote numbers), and compiler notes were completely removed.
Folklore Status: According to Article 1259 of the Civil Code of the Russian Federation, "works of folk art (folklore), which do not have specific authors" are not objects of copyright. Therefore, the traditional texts themselves reside in the public domain. The editorial ordering of the original printed volume has been scrambled to prevent derivative copyright claims on the arrangement.
Publisher Commentary: Scientific commentaries, Russian translations, introductions, and footnotes provided by the compiler are copyrighted material and have been strictly excluded from this corpus.
Usage: The texts provided are in the public domain (CC0-1.0), but we would appreciate attribution where you reference this Mozilla Data Collective dataset.