Release Date: 8/5/2026
Format: TXT
Size: 12.38 KB
Share
This corpus contains a collection of traditional Chechen proverbs and sayings (кицанаш) digitized from I. Yu. Aliroev's 1990 bilingual collection "Кувшин мудростей" (A Jug of Wisdom), published in Grozny. The dataset provides clean data designed to catalyze NLP research for the lesser-resourced Chechen language, a Nakh language of the Northeast Caucasus written in an extended Cyrillic alphabet. Because proverbs and sayings are rich in cultural realias, historical context, and idiomatic phrasing, this corpus is uniquely positioned to serve as a foundation for "cultural competence" benchmarks, allowing researchers to evaluate an LLM's understanding of language-specific cultural nuances (e.g., through cloze tests or masked word evaluations). To ensure strict adherence to public domain copyright laws, the source's Russian translations and all editorial apparatus have been excluded, retaining only the public-domain Chechen folklore. The palochka is uniformly encoded as U+04C0.
Restrictions/Special Constraints
None
Forbidden Usage
None
This corpus contains a collection of traditional Chechen proverbs and sayings (кицанаш). Chechen is a Nakh language of the Northeast Caucasian (Nakh-Daghestanian) family, spoken primarily by the Chechen people in the Chechen Republic (Russia). As a lesser-resourced language written in an extended Cyrillic alphabet, digitized native texts are relatively scarce. The texts in this corpus were digitized from I. Yu. Aliroev's 1990 bilingual (Chechen-Russian) collection.
Beyond cultural preservation, this dataset is specifically designed for evaluating the "cultural competence" of Large Language Models (LLMs) in the Chechen language. Proverbs and sayings contain dense cultural realias, idiomatic structures, and language-specific semantic dependencies. This corpus can be used to build fill-in-the-blank benchmarks (where specific words are omitted) to test whether an LLM actually understands regional cultural nuances rather than just generating statistically probable, yet culturally empty, translations.
Chechen orthography uses the palochka (Ӏ, U+04C0) as a full letter: on its own it marks a glottal stop / pharyngeal, and it forms the digraphs гӀ, кӀ, пӀ, тӀ, хӀ, цӀ, чӀ. In this corpus the palochka is always encoded as U+04C0 (uppercase) — never as the Latin letters "I"/"l" or the digit "1", which scanned and OCR'd sources frequently substitute for it. Consumers normalizing or tokenizing the text should treat U+04C0 as a letter, and may want to guard against those look-alike substitutions when merging in data from other sources.
Total Word Count: 2927
Files:
proverbs.txt: 2927 words (409 entries)
The following volume served as the source material for this corpus:
Aliroev, I. Yu. (1990). Кувшин мудростей: Чеченские пословицы и поговорки [A Jug of Wisdom: Chechen Proverbs and Sayings]. Grozny: Checheno-Ingush "Kniga" Publishing.
The dataset is provided in plain text. The file begins with a single YAML Front Matter block containing the compilation metadata. Following this header, the entries are listed sequentially, with each entry delimited by a blank line.
Below is a verbatim segment of the beginning of the file:
---
title: "Кувшин мудростей: Чеченские пословицы и поговорки"
lang: "ce"
year: "Unknown"
source_original: "Unknown"
source_container: "Алироев, И. Ю. (1990). Кувшин мудростей: Чеченские пословицы и поговорки. Грозный: Чечено-Ингушское издательско-полиграфическое объединение «Книга»."
---
КУВШИН МУДРОСТЕЙ: ЧЕЧЕНСКИЕ ПОСЛОВИЦЫ И ПОГОВОРКИ
Аьрга стом санна муьста ма хила — церг тоьх-тоьхначо дӀатосур ву хьо.
Аьхка лаьхьанах кхеравелларг — Ӏай карсанах кхеравелла.
Аьхка мало, Ӏай хало.
Аьхка Ӏиллинарг, Ӏай идда.
Аьхкенан заманчохь мукъа леллачунна, Ӏаьно тӀе ког боккху.
Note: The entries follow the source volume's plain alphabetical order (by Chechen initial letter). A straight alphabetical listing is a mechanical, non-original arrangement, so no counterpart to the compilers' arrangement is reproduced or needs to be neutralized; the alphabetic section headings themselves are not reproduced.
This dataset was created through the following pipeline:
Acquisition: Sourced a scanned PDF of the 1990 volume, in which each Chechen proverb is paired with a Russian translation.
Transcription: The source pages were transcribed by a vision-language model (Gemini 3.1 Pro), prompted to return only the Chechen proverbs. Unlike a fixed OCR engine, the model reads the extended-Cyrillic Chechen letters — including the palochka (Ӏ) — directly, rather than substituting look-alike glyphs such as "I", "l", or "1".
Exclusion of the Copyrighted Layer: Only the public-domain Chechen folklore was retained; the compiler/translator's Russian translations, the introduction, and all other editorial apparatus were excluded.
Proofreading: Every proverb was then proofread by hand against the source PDF, line by line, and corrected wherever the transcription diverged from the printed text. The palochka is uniformly encoded as U+04C0 throughout.
Folklore Status: According to Article 1259 of the Civil Code of the Russian Federation, "works of folk art (folklore), which do not have specific authors" are not objects of copyright. Therefore, the traditional proverb texts themselves reside in the public domain.
Publisher Commentary and Translations: The compiler's Russian translations, introduction, scientific commentary, and any footnotes are copyrighted material and have been strictly excluded from this corpus; only the Chechen folklore texts are included.
Arrangement: The entries are presented in the source's plain alphabetical order and the alphabetic section headings are omitted. A straight alphabetical ordering is not an original compilation arrangement, so no editorial arrangement of the source is reproduced.
Usage: The texts provided are in the public domain (CC0-1.0), but we would appreciate attribution where you reference this Mozilla Data Collective dataset.