Release Date: 8/5/2026
Format: TXT, TSV
Size: 132.39 KB
Share
This corpus contains a collection of traditional Bulgarian proverbs and sayings (пословици и поговорки) digitized from the 1986 third edition of the compilation by M. Grigorov and K. Katsarov, published by Nauka i izkustvo in Sofia. The dataset provides clean data designed to catalyze NLP research for the Bulgarian language. Because proverbs are rich in cultural realias, historical context, and idiomatic phrasing, this corpus is uniquely positioned to serve as a foundation for "cultural competence" benchmarks, allowing researchers to evaluate an LLM's understanding of language-specific cultural nuances (e.g., through cloze tests or masked word evaluations). The compilers' editorial apparatus and thematic section headings have been excluded and the entries are presented in plain alphabetical order. Every entry was additionally checked against two independent published collections; the per-entry result ships with the dataset as corroboration.tsv, with 55.8% of entries having a close counterpart and a further 19.2% a partial one.
Restrictions/Special Constraints
None
Forbidden Usage
None
This corpus contains a collection of traditional Bulgarian proverbs and sayings (пословици и поговорки). The texts were digitized from the 1986 third revised edition of the compilation by M. Grigorov and K. Katsarov, published by Nauka i izkustvo in Sofia.
Beyond cultural preservation, this dataset is specifically designed for evaluating the "cultural competence" of Large Language Models (LLMs) in the Bulgarian language. Proverbs contain dense cultural realias, idiomatic structures, and language-specific semantic dependencies. This corpus can be used to build fill-in-the-blank benchmarks (where specific words are omitted) to test whether an LLM actually understands regional cultural nuances rather than just generating statistically probable, yet culturally empty, translations.
Total Word Count: 40262
Files:
proverbs.txt: 40262 words (6142 entries)
corroboration.tsv: per-entry attestation report (6142 rows)
The following volume served as the source material for this corpus:
Grigorov, M., & Katsarov, K. (Comps.). (1986). Български пословици и поговорки [Bulgarian Proverbs and Sayings] (3rd rev. ed., print run 74,102). Sofia: Nauka i izkustvo.
The digital file used is mirrored on the Internet Archive as item
20231013_20231013_0636.
The following independent collections were used to corroborate the entries. No text from them is included in this corpus:
Minkov, Ts. (Ed.). (1963). Българско народно творчество в дванадесет тома: Т. 12. Пословици, поговорки, гатанки [Bulgarian Folk Art in Twelve Volumes: Vol. 12. Proverbs, Sayings, Riddles]. Sofia: Bulgarski pisatel.
Slaveykov, P. R. (2003). Български притчи или пословици и характерни думи [Bulgarian Parables or Proverbs and Characteristic Words]. Sofia: Zahariy Stoyanov (Българска класика series). Originally published 1889–1897; the proverb text is public domain.
The dataset is provided in plain text. The file begins with a single YAML Front Matter block containing the compilation metadata. Following this header, the entries are listed sequentially, with each entry delimited by a blank line.
Below is a verbatim segment of the beginning of the file:
---
title: "Български пословици и поговорки"
lang: "bg"
year: "Unknown"
source_original: "Unknown"
source_container: "Grigorov, M., & Katsarov, K. (Comp.). (1986). Български пословици и поговорки [Bulgarian Proverbs and Sayings]. Sofia: Nauka i izkustvo."
---
БЪЛГАРСКИ ПОСЛОВИЦИ И ПОГОВОРКИ
Абе то ще се мре, ами току здраве да е.
Авлигата е малка, а кога брани гнездото си, на усойницата надвива.
Агне се в чувал не купува.
Агнешките кожи са по-много на пазар от овчите.
Адет не е закон, ама закон става.
Note: The compilers' 209 thematic section headings are not reproduced, and the entries are ordered alphabetically rather than in the arrangement of the printed volume.
corroboration.tsv records, for every entry, whether a counterpart was
found in an independent published collection. Columns: n (position in
proverbs.txt), score (0–1 similarity to the closest reference line),
status, and attested_in (which collection supplied the match).
| Status | Meaning | Count | Share |
|---|---|---|---|
| ATTESTED | close counterpart (score >= 0.80) | 3429 | 55.8% |
| PARTIAL | partial counterpart (0.62–0.80) | 1180 | 19.2% |
| NONE | no counterpart found | 1533 | 25.0% |
Matching normalizes orthographic and dialectal variation (stress marks, pre-1945 letterforms, са/се alternation) before comparison, so ATTESTED includes spelling variants of the same entry rather than exact matches only. Candidate lines were drawn from the full text of each reference volume, so a small number of matches may fall in an editor's apparatus that quotes an entry rather than in the entry list proper.
A NONE result does not imply the entry is erroneous. The reference collections are independent works, not supersets of the source volume, and they differ in regional coverage. The report is provided so that users can weight or filter entries by attestation according to their own needs; for benchmark construction, restricting to ATTESTED entries yields a conservative subset.
This dataset was created through the following pipeline:
Acquisition: The source was a third-party digital transcription of the 1986 volume, distributed as a PDF exported from Microsoft Word (document metadata attributes it to spiralata.net). It is not a scan of the printed book; the Internet Archive copy cited above is a rendering of the same file and shares its defects.
Text Extraction: The PDF text layer was extracted with
pdftotext -layout; line-wrapped entries were rejoined, page furniture
and thematic headings removed.
Correction: The extracted text was proofread against the source PDF. Roughly 20 defects were corrected, including character substitutions inherited from the upstream transcription, lost hyphens in comparative forms, Latin characters used as stress marks inside Cyrillic words, entries split across page breaks, and spacing errors. Two irrecoverably truncated entries were removed. Archaic and dialectal forms were retained as printed and not modernized.
Corroboration: Every entry was matched against two independent
published collections (Minkov 1963; Slaveykov 1889–1897) after
normalization, and the per-entry result recorded in
corroboration.tsv.
Because the upstream source is a transcription rather than a scan, defects it introduced cannot be detected by comparing the output against it. The correction pass in step 3 was manual and is not assumed exhaustive; the corroboration report in step 4 exists to give users an independent signal per entry. Two readings remain unresolved and are retained as printed in the source.
Folklore Status: Under Art. 4(3) of the Bulgarian Law on Copyright and Neighbouring Rights (ЗАПСП), фолклорни творби (works of folklore) are not objects of copyright. The proverb texts themselves are therefore in the public domain.
Compilation: Art. 11 ЗАПСП vests rights in compilations in the person who performed the selection or the arrangement. The compilers' arrangement — the thematic grouping and the order of entries within it — is not reproduced here: headings are omitted and entries are sorted alphabetically.
Publisher Commentary:
The editor's note, introduction, thematic index, and the compilers'
bracketed explanatory notes are copyrighted material and have been
excluded from this corpus. No text from the two corroboration volumes is
included; corroboration.tsv records only similarity scores and which
collection matched.
Usage: The texts provided are in the public domain (CC0-1.0), but we would appreciate attribution where you reference this Mozilla Data Collective dataset.