License:
CC-BY-NC-4.0
Steward:
CommunityDataset ID:
cmuvj8h64005ko107pwsqx7jj
Task: N/A
Release Date: 10/5/2026
Format: CSV, JSONL, TXT, MD
Size: 577.46 KB
The Kanuri Lexicon is a structured lexical dataset derived from the A–Z Kanuri Dictionary attributed to Professor Bulakarima. It contains Kanuri headwords, grammatical labels, English definitions and source-entry information covering a broad range of Kanuri vocabulary. The resource has been structured from the original dictionary for use in language documentation, lexical research, natural language processing, language education and the development of Kanuri language technologies. The dataset preserves the source terminology and does not claim to provide additional linguistic annotation beyond what is present in the source dictionary.
Licensing
Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)
https://spdx.org/licenses/CC-BY-NC-4.0.htmlRestrictions/Special Constraints
This dataset is provided for research, educational, language documentation, linguistic analysis, natural language processing and language technology development. Use of the dataset is subject to the selected Creative Commons Attribution-NonCommercial 4.0 International licence and the additional terms stated in this datasheet. Users should preserve attribution to Professor Bulakarima as the copyright holder and acknowledge CIATECH Africa as the technical curator when redistributing or publishing substantial portions of the dataset. Users should not represent the dataset as an independently created work or as an authoritative linguistic standard without appropriate scholarly verification.
Forbidden Usage
The dataset must not be used for commercial purposes where such use is prohibited by the selected CC BY-NC-4.0 licence. Users must not remove or obscure copyright, attribution or provenance information contained in the dataset. Users must not falsely represent the dataset, its entries or its derived resources as being independently authored or officially endorsed by CIATECH Africa. Users should not use the dataset in a manner that deliberately misrepresents Kanuri language, culture or communities.
Ethical Review
This dataset is derived from an existing published/compiled dictionary source rather than from a new collection of human participant data. No new human subjects were recruited or recorded as part of the preparation of this dataset, and the dataset does not contain speaker recordings or participant-level personal data. The primary ethical consideration is intellectual property and authorization to reproduce and distribute the source material. Professor Bulakarima is identified as the copyright holder, and public distribution should proceed only after appropriate authorization and confirmation of the applicable licence. CIATECH Africa has applied source-preserving technical curation and has not intentionally added personal or sensitive participant information.
Intended Use
This dataset is intended for Kanuri language documentation, lexical research, linguistic analysis, natural language processing, machine learning research, language modelling, machine translation research, spellchecking, information retrieval, language education and the development of Kanuri language technologies. It may also serve as a foundational text resource for future Kanuri AI systems when used in accordance with the applicable licence and attribution requirements.
Source: KANURI DICTIONARY FROM LETTER A TO Z. Source format: DOCX. Source length: 301 pages. Structured records: 9,554. Unique headwords: 9,383. The dataset was programmatically structured from the source document while preserving the source lexical content. Unicode text was normalized and whitespace was standardized for machine readability. Exact duplicate records were removed during structuring. No external corpus, machine-generated translation, speech data or model-generated content was added. The dataset includes lexical entries with grammatical labels and English definitions. The source contains multiple grammatical categories, including nouns, verbs, verbal nouns, adjectives, adverbs, ideophones, pronouns, particles and other categories, as well as cross-references and examples. This version should be considered a source-derived lexical dataset (v1.0.0) and not a linguistically validated gold-standard corpus. Researchers should independently validate entries before using the resource for high-stakes linguistic modelling or publication. The dataset forms part of CIATECH Africa's broader effort to develop digital language resources for Kanuri and the wider Lake Chad region.