Release Date: 9/2/2026
Format: DOCX, TXT
Size: 74.83 KB
Share
The Wakhi Literature Corpus is a collection of literary texts in the Wakhi language, representing the linguistic, literary, and cultural heritage of the Wakhi-speaking community. Wakhi is an Eastern Iranian language spoken primarily in the mountainous regions of northern Pakistan, as well as parts of Afghanistan, Tajikistan, and China. As a comparatively low-resource language, Wakhi has limited digitized linguistic and computational resources, making this corpus valuable for language documentation and preservation. The corpus brings together Wakhi literary material that can support the study of the language's vocabulary, grammar, writing practices, literary traditions, and cultural expression. Literary texts provide researchers with naturally occurring examples of language use and offer insights into how Wakhi is used to express narratives, ideas, experiences, traditions, and cultural knowledge.
Licensing
Creative Commons Attribution No Derivatives 4.0 International (CC-BY-ND-4.0)
https://spdx.org/licenses/CC-BY-ND-4.0.htmlRestrictions/Special Constraints
Use of the dataset should comply with applicable copyright, licensing, and attribution requirements.
Forbidden Usage
Do not use or redistribute the dataset in ways that violate applicable licenses, copyrights, or rights of the original creators.
Intended Use
Intended for Wakhi language documentation, linguistic and literary research, education, and cultural preservation
The Wakhi Literature Corpus is a collection of literary texts in the Wakhi language, representing the linguistic, literary, and cultural heritage of the Wakhi-speaking community. The corpus provides digitized literary material that can support language documentation, preservation, linguistic research, literary studies, and the development of resources for low-resource language technologies.
Wakhi is an Eastern Iranian language spoken primarily in the mountainous regions of northern Pakistan, Afghanistan, Tajikistan, and western China. It is a relatively low-resource language with limited digitized textual and computational resources.
Language: Wakhi
ISO 639-3: wbl
a, b, č, č̣, d, ḍ, ð, dz, e, ə, f, g, ɣ, ɣ̌, h, i, j, ǰ, k, l, m, n, o, p, q, r, s, š, ş, t, ţ, θ, ts, u, ʉ, v, w, x, x̌, y, z, ž, ẓ
The dataset contains the following file formats:
DOCX — Microsoft Word document
TXT — Plain text document
Wakhi/
├── 1. Khéṭk jumlayisht.docx
└── wkhi.txt
S̃hẽru cẽ k̃hẽdhoyẽ nungẽn carẽm. Khẽdoyi ramdil woz miribon. Khẽdhoy sakẽ mẽdad wost. Amnẽt amoni k̃hẽdhoyẽr khus̃h. Khẽdhoyri amnẽt amon khus̃h. Khẽdhoyẽ k̃het dẽnyoẽ, yawẽ k̃het dẽsrũrẽ sokht lecrit. Khẽdhoy sakẽr baf k̃hakẽ tofiqẽ rand. Dẽnyoẽ, Khẽdhoyẽ k̃het dẽstũrẽ sokht lecẽrit. Shak me gok̃hit Khẽdhoy dẽstũrẽb bedht. Khẽdhoy k̃hat haramẽ pẽzũv kart, kuyẽs̃h ki shak gok̃t. Khẽdhoyẽ ki ramẽ pẽzũv ne kart beti chora nast. Chorayi nast, Khẽdhy ramẽ pẽzũv kart. Amnẽt amoni dẽlemanẽn helakẽr dẽrkor. Amnẽt amon ki ne wost dẽlẽmanẽb ne hẽlakẽ bas wezin. Cẽrg amnẽt amonẽn bẽshkhab dẽlẽmanẽn halẽna? Amnẽs̃h chizẽr k̃hanẽ? Dẽ sẽpo mulkẽb amn cersokht wizit? Shak yarkẽr pẽtmũyni baf nat. Yash khalgvẽ pũtmuyd.