License:
NOODL-1.0
Steward:
Institute of African Digital HumanitiesDataset ID:
cmu1b093003ulnz07ewduih9f
Release Date: 9/14/2026
Format: TSV, PDF
Size: 79.64 KB
A Kotoko Sociocultural Dataset is a sociocultural and onomastic dataset documenting the Kotoko community of Cameroon (a sultanate people (self-named Mantagé) claiming descent from the ancient Sao civilisation), collected under the project "Projet de confection d'un didacticiel sur la composition des peuplements au Cameroun" (a didactic-software project on the composition of Cameroon's settled peoples), coordinated by Prof. Emmanuel Ngue Um for CERDOTOLA (the Inter-State Institution of scientific co-operation for the preservation, diffusion and valorization of the African cultural heritage) during the project's pilot data-collection phase in March-April 2020. Unlike a purely linguistic corpus, the dataset foregrounds the community itself as the unit of documentation: its history, social organisation, religious life and material culture, with onomastics, common expressions and a numeral system providing a linguistic component elicited in the community's own speech (recorded in the International Phonetic Alphabet, or in a Latin-based approximation where IPA was not available to the fieldworker). Self-designated Mantagé and known to neighbours as Kotoko, the group is today so thoroughly integrated with the Massa and Mousgoum that the three are difficult to distinguish; its speech is also found at Magba and Lagdo. Oral tradition traces an ancient Egyptian origin, with intermediate stays in Sudan and Chad, and the group considers itself the founding and first-settling population of its present town. The group is led by a sultan at the head of a first-degree sultanate, advised by notables and three canton chiefs who must be consulted before any decision is made public; descent is patrilineal. Marriage proceeds from a formal hand-request to an agreed engagement date, sealed with kola, the *Néné* and a cash gift. A cult addressed to water spirits secures safe and abundant fishing, and the appearance of a monitor lizard is read as an omen for the sultanate. No specific taboos were reported. Men fish and farm, women grow market vegetables; okra (*gombo*), *lalo* and *bocco* are staple dishes, and tea and porridge everyday drinks. Women wear the tailored *kaba* dress and jewellery for dancing; the group's Koungouli, Kolo and Daramya rhythms and its claim of Sao descent are shared with neighbouring Massa and Mousgoum. The dataset's primary added value lies in documenting the Kotoko community's own self-description — its social organisation, kinship and marriage practices, religious and symbolic life, economy and material culture, and internal system of personal, ethnic and place names — rather than only its named language or dialect. This community-centred, ethnographic design distinguishes it from language-only documentation, while the accompanying onomastic, common-expression and numeral-system data still make the dataset usable for lexicographic and comparative linguistic work on the Kotoko variety.
Licensing
Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)
https://licensingafricandatasets.com/nwulite-obodo-licenseRestrictions/Special Constraints
By downloading this dataset, you agree: - To use it for research and scientific use only - that you will not re-host or re-share this dataset
Forbidden Usage
You agree not to use the data for: profiling, surveillance or any purpose prejudicial to the interests of the individuals or community described; caricaturing, stereotyping, or otherwise misrepresenting the community documented; Generative AI; reproduction; duplication; modification; augmentation; copying; distribution; transmission; display; sale; transfer; publication or creation of derivative works without the explicit permission of the legal owner of the dataset.
Intended Use
(a) Ethnographic and cultural documentation: - Sociocultural and anthropological research: The narrative responses (identity, residence and mobility, social and political organisation, kinship and marriage, religious and symbolic life, economy and material culture) support ethnographic, anthropological and sociolinguistic research on the Kotoko community, and complement rather than replace existing ethnographic and anthropological scholarship. (b) Onomastic and toponymic research: - Personal-name, ethnonym and toponym documentation: The structured onomastic entries, each tagged with a semantic primitive and grammatical category, support onomastic and lexicographic research and comparative naming-practice studies across Cameroonian communities. (c) Linguistic and lexicographic tasks: - Lexicon documentation and comparative numeral typology: The common-expression and numeral-system tables, transcribed where possible in IPA, are useful for building basic lexical resources for the Kotoko variety and for comparative numeral-system typology across Cameroonian languages. (d) Pedagogical and educational-software use: - Didactic-software content: As part of the pilot phase of a project building educational (“didacticiel”) material on the composition of Cameroon's settled peoples, the dataset supports the development of teaching content on the Kotoko community's history, culture and onomastics. (e) Cultural modeling and intercultural dialogue: - Cultural modeling: The sociocultural, onomastic and linguistic data in this dataset can be used to build and enrich interactive network-based representations of cultural communities, modeling the relationships and affinities between Cameroonian communities across historical, linguistic, geographic, social and cognitive dimensions. An example of such cultural modeling in the African multicultural and multilingual context is the CamRhizome interactive visualization, accessible at https://ngue-um.github.io/Enseignement-Apprentissage-LN/#/reseau, which maps intercommunal connections among 20 Cameroonian language communities and can serve as a tool for enhancing intercultural dialogue in African multicultural and multilingual spaces.
This dataset is distributed as a .tar.gz archive (A_Kotoko_Sociocultural_Dataset.tar.gz) with the following contents:
A_Kotoko_Sociocultural_Dataset/
├── 1_Metadonnees_Enquete.tsv — administrative metadata (1 row)
├── 2_Reponses_Ouvertes.tsv — open-ended narrative Q&A responses
├── 3_Onomastique.tsv — onomastic entries (anthroponyms, ethnonyms, toponyms)
├── 4_Expressions_Usuelles.tsv — common bilingual expressions
├── 5_Systeme_Numeral.tsv — numeral system entries
├── 6_Annexe_Primitifs.tsv — NSM semantic primitives reference table (65 items)
└── A_Kotoko_Sociocultural_Dataset.pdf — PDF version of the questionnaire
Note: The fieldwork interview audio recording is not yet included in this release. It will be added in a subsequent update. Recordings were collected on the fieldwork sites and document the interaction between the interviewer and the interviewee.
Self-designated Mantagé and known to neighbours as Kotoko, the group is today so thoroughly integrated with the Massa and Mousgoum that the three are difficult to distinguish; its speech is also found at Magba and Lagdo. Oral tradition traces an ancient Egyptian origin, with intermediate stays in Sudan and Chad, and the group considers itself the founding and first-settling population of its present town.
The group is led by a sultan at the head of a first-degree sultanate, advised by notables and three canton chiefs who must be consulted before any decision is made public; descent is patrilineal. Marriage proceeds from a formal hand-request to an agreed engagement date, sealed with kola, the Néné and a cash gift.
A cult addressed to water spirits secures safe and abundant fishing, and the appearance of a monitor lizard is read as an omen for the sultanate. No specific taboos were reported.
Men fish and farm, women grow market vegetables; okra (gombo), lalo and bocco are staple dishes, and tea and porridge everyday drinks. Women wear the tailored kaba dress and jewellery for dancing; the group's Koungouli, Kolo and Daramya rhythms and its claim of Sao descent are shared with neighbouring Massa and Mousgoum.
The interview was conducted with the local expert speaking Kotoko. Rubrique 4 of the questionnaire elicited 157 common bilingual expressions (salutations, politeness formulas and everyday phrases), of which 146 have an endogenous-language form recorded to date; rubrique 5 elicited a 107-item numeral system, of which 107 have an endogenous-language form recorded to date. Together with the anthroponyms, ethnonyms and toponyms in rubrique 3, these fields constitute the dataset's linguistic component and are transcribed, where possible, in the International Phonetic Alphabet (IPA/API), per the questionnaire's instructions; where a fieldworker was not proficient in IPA, a Latin-alphabet approximation was used instead. No audio recordings are included in this release, although the questionnaire's protocol called for the full interview to be captured on dictaphone.
The dataset was collected through a questionnaire designed to gather sociocultural and onomastic information about the Kotoko community. This was done during the pilot data-collection phase (March-April 2020) of the "Projet de confection d'un didacticiel sur la composition des peuplements au Cameroun," coordinated by Prof. Emmanuel Ngue Um for CERDOTOLA (Centre International des Civilisations Bantu / Institution Inter-Etats de cooperation scientifique pour la preservation, la diffusion et la mise en valeur du patrimoine africain), in partnership with the Ecole normale superieure de Bertoua. Record identifier: KTKK_2020_001.
The dataset represents a sociocultural and onomastic questionnaire designed to elicit the identity, history, social organisation, religious life, economy and material culture of a named community in Cameroon, together with a structured record of its personal names, ethnic names and place names, common bilingual expressions, and numeral system. The questionnaire was developed for the pilot phase of a project building educational (“didacticiel”) material on the composition of Cameroon's settled peoples; unlike a purely linguistic questionnaire, its entry point and primary unit of documentation is the community, not its language or dialect, although a linguistic component (onomastics, common expressions, numeral system) is elicited throughout in the community's own speech.
The dataset comprises 51 narrative question/answer entries (rubriques 2 and 3.1-3.6); 21 anthroponym entries, 0 ethnonym entries and 3 toponym entries (rubrique 3); 146 of 157 common expressions with an endogenous-language form recorded (rubrique 4); 107 of 107 numeral-system entries with an endogenous-language form recorded (rubrique 5); and a fixed 65-item reference table of NSM semantic primitives (annex).
This release comprises six related tables, distributed here as separate tab-separated-value (TSV) files (one per original spreadsheet sheet) alongside the PDF version of this submission form:
1_Metadonnees_Enquete.tsv: One row of administrative metadata for the questionnaire: record identifier, questionnaire title/version/author, the group's administrative and emic (self-given) names and their meaning, names given by/to neighbouring groups, location, fieldwork date, language(s) used, and consultant information.
2_Reponses_Ouvertes.tsv: Question/answer pairs for the open-ended narrative sub-sections of rubriques 2 (sociocultural profile: identity, residence and mobility, social and political organisation and kinship, religious life, socio-economic and sociocultural life) and 3.1.1-3.1.6 (general anthroponymic structure), with the original French question, the free-text answer, and the language(s) in which the answer was given.
3_Onomastique.tsv: One row per onomastic entry (anthroponym, ethnonym or toponym): orthographic form, IPA transcription where available, literal meaning/etymology, an associated NSM semantic primitive and grammatical category, and (for anthroponyms) the type of name.
4_Expressions_Usuelles.tsv: One row per common bilingual expression (greetings, politeness formulas and everyday phrases): the French prompt, the endogenous-language form where recorded, an IPA transcription where available, and a word-for-word gloss.
5_Systeme_Numeral.tsv: One row per numeral (1 upward): the French label, the endogenous-language form where recorded, an IPA transcription where available, and a compositional gloss (e.g. 7 = 6+1).
6_Annexe_Primitifs.tsv: The fixed reference table of Natural Semantic Metalanguage (NSM) primitives (Goddard & Wierzbicka) used to tag onomastic entries in sheet 3.
Anthroponyms (from 3_Onomastique.tsv):
| Anthroponym | Meaning / notes |
|---|---|
| Barounga [ baruba ] | Celui qui vient après les jumeaux |
| Hilinjwa [ hiligwa ] | Jumeaux |
| Hassana [ hasana ] | Le premier des jumeaux garçon |
| Hissein [ hisen ] | Le deuxieme né |
| Haoua [ haua ] | La 1ere née jumeaux filles |
Common expressions (from 4_Expressions_Usuelles.tsv):
| French | Kotoko | Word-for-word gloss |
|---|---|---|
| Salut (formule neutre, à toute heure) | Baraku | Salut |
| salutations du matin (une personne) | Okizu | Bonjour |
| salutations du matin (plusieurs personnes) | Ure kizu | Vous demain |
| salutations du matin (en signe de déférence/respect) | Ure kizu | Vous demain |
| salutations de la mi-journée (une personne) | Ka wéda | Comment tu as la journée |
Numeral system (from 5_Systeme_Numeral.tsv):
| French | Kotoko |
|---|---|
| Un | Sekeɗi |
| Deux | Kitcho |
| Trois | Kaker |
| Quatre | kadé |
| Cinq | Chéchi |
| Six | Vrekaker |