License:
NOODL-1.0
Steward:
Institute of African Digital HumanitiesDataset ID:
cmruvl28e00b4md078sjgbjo9
Release Date: 7/21/2026
Format: MP3, TSV
Size: 7.56 MB
Share
Babute-Njore_ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and technological enhancement of Babouté (also known as Vute, Bute, Wute, Voute; ISO 639-3: vut), a Bantoid language spoken by farming communities across the Centre, Adamawa and East Regions of Cameroon. This release documents the variety of Babouté spoken in Njoré, a Vute-speaking locality in the Mbandjock area (Haute-Sanaga division, Centre Region). Babouté remains sparsely represented in computational language resources despite an estimated 21,000 speakers and a documented standard orthography. The dataset comprises three closely aligned components: (i) a datasheet containing lexical entries and example sentences reflecting attested usage in Babouté as spoken in Njoré; (ii) high-quality audio recordings of a substantial subset of these entries, produced by a native speaker; and (iii) explicit audio-sentence mapping files enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit documentation of the Njoré variety of Babouté, a Bantoid language of central Cameroon that — despite a comparatively larger speaker base and an existing General Alphabet of Cameroonian Languages (GACEL)-based orthography (standardized in 1979) — remains, like many regional languages of Cameroon, poorly represented in modern computational and pedagogical resources at the level of specific local varieties. The parallel availability of text in Babouté and in French, together with aligned speech for a substantial subset of entries, makes the dataset suitable for a range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment and pronunciation modelling. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The datasheet's word-for-word parsing of both the Babouté and French example sentences further supports morphological analysis and glossed-corpus studies. At the same time, the structured datasheet supports basic lexicographic and grammatical documentation, and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the Babute-Njore_ALCAM-MultimodalDataset exemplifies an approach to Cameroonian language resources that documents linguistic variation at the level of specific villages and speech communities, rather than only at the level of the named language as a whole.
Licensing
Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)
https://licensingafricandatasets.com/nwulite-obodo-licenseRestrictions/Special Constraints
By downloading this dataset, you agree: - To use it for research and scientific use only - that you will not re-host or re-share this dataset
Forbidden Usage
You agree not to use the data for: determining the identity of the speaker(s) in the dataset; attempt to clone the voice or train models that imitate the speaker(s) in this dataset; Generative AI; reproduction; duplication; modification; augmentation; copying; distribution; transmission; display; sale; transfer; publication or creation of derivative works without the explicit permission of the legal owner of the dataset.
Intended Use
(a) Speech-related tasks: - Automatic speech recognition (ASR): Audio-text alignment allows the evaluation of speech recognition models for Babouté. It should be noted that the sentences are transcribed using the IPA alphabet, rather than the GACEL-based practical orthography that has existed for the language since 1979. - Text-to-speech (TTS): As the dataset contains clean sentence-audio pairs, it can also be used to evaluate speech synthesis or text-to-speech models. The use of IPA transcription rather than the standardized practical orthography should be taken into account when designing TTS experiments. - Speech-text alignment/forced alignment benchmarking: Fine-grained audio-text pairing provides ground truth for evaluating phoneme- or word-level aligners adapted to tonal Bantoid languages of Cameroon. (b) Translation and multilingual tasks: - Machine translation (Babouté ↔ French): The sentence-level alignment between Babouté (`LangEx`) and French (`FrenchEx`) makes the datasheet usable as a small parallel corpus for evaluating translation models, with the caveat that the Babouté side uses IPA transcription rather than the standardized orthography, and that aligned audio is only confirmed for a subset of entries. (c) Linguistic and lexicographic tasks: - Morphological analysis/glossed-corpus studies: The word-for-word parsing columns (`LangPars`, `FrenchPars`) support computational morphology, interlinear text (ILT) modelling, and grammar induction for Babouté and other Bantoid languages of Cameroon. - Lexicon documentation and part-of-speech tagging: The `Word`, `French` and `POS` columns are useful for building basic lexical resources and part-of-speech taggers for Babouté as spoken in Njoré, a locally documented variety that remains largely absent from computational resources. - Language documentation and revitalisation: The aligned audio and text support pedagogical uses in teacher training and community-based language revitalisation efforts focused on the Njoré speech community.
Babouté (also known as Vute, Bamboute, Boute, Bute, Foute, Pute, Voute, Voutere, Woute, Wute; ISO 639-3: vut) is a Bantoid language (Niger-Congo > Atlantic-Congo > Benue-Congo > Mambiloid > Mambila-Vute > Vute) spoken by roughly 21,000 people (1997 estimate) mainly in Cameroon, with some speakers reported in Nigeria. Speakers are distributed across three areas: the Mayo-Banyo and Djérem divisions of the Adamawa Region (near Banyo and Tibati); the Mbam-et-Kim, Mbam-et-Inoubou and Haute-Sanaga divisions of the Centre Region (near Ngoro, Nanga-Eboko and Mbandjock); and the Lom-et-Djérem division of the East Region (near Doumé). This dataset documents the variety spoken in Njoré, a Vute-speaking locality in the Mbandjock area of the Haute-Sanaga division, Centre Region, founded historically by settlers associated with the broader Vuté migration across the Sanaga river.
Published descriptions of Babouté (Vute) distinguish three broader dialect clusters — Eastern, Central and Doumé — that differ in vowel/diphthong inventory and in the preservation of certain labialized consonants. This release documents a single, localized variety as elicited from a native speaker in Njoré; no formal correspondence between the Njoré variety and the Eastern/Central/Doumé dialect clusters described in the literature has been established here, and this should be clarified with the dataset's creator before downstream use. Unlike the multi-dialect structure adopted for some other languages in the same ALCAM series (e.g. Batanga, which documents the distinct Banoho and Bapuku varieties separately), no further dialect-internal split is documented within this dataset.
The writing system used for the transcription of Babouté in this dataset is the International Phonetic Alphabet (IPA), as reflected in the Word, LangEx and LangPars columns of the datasheet and in the sentence field of the audio-sentence mapping files. The phonological inventory below draws on the 374 datasheet entries and the 245 unique audio-aligned sentences (the mapping.tsv files across the dataset's four recording batches). A standardized practical orthography for Babouté (Vute), based on the General Alphabet of Cameroonian Languages (GACEL), has existed since 1979, but is not the transcription system used in this dataset.
The data attests seven oral vowel qualities: a, e, ə, i, o, ɔ, u. Note that this transcription uses the central vowel ə (schwa) where some published descriptions of Babouté (Vute) instead report ɨ or ɛ; no nasalized vowels or diphthongs are systematically marked in the transcribed data, in contrast to the fuller vowel inventory (including nasal vowels and diphthongs such as /ei/, /ai/, /ɨi/, /əi/, /oi/) reported for Vute in the wider literature. A combining tilde-below diacritic (◌̰), which in IPA conventionally marks creaky voice, occurs on 135 tokens across the Word, LangEx and LangPars columns; its precise phonetic conditioning has not been established here and should be clarified with the dataset's creator before downstream use.
The consonant inventory attested in the datasheet and mapping files includes:
b, d, g, j, k, l, m, n, p, s, t, w, y — the core set of simple consonants, plus two implosives, ɓ (voiced bilabial implosive, 138 tokens) and ɗ (voiced alveolar implosive, 60 tokens), alongside the nasals ŋ (96 tokens) and ɲ (20 tokens).
A glottal stop ʔ (23 tokens) is also attested, alongside an apostrophe (') that recurs independently across 56 tokens in the transcribed sentences; as with the glottal stop, its precise grammatical or phonetic function relative to ʔ has not been formally documented and should be clarified with the dataset's creator.
Complex onsets and prenasalized consonants are attested: mb, nd, ng (ŋg), nj, kp, gb, matching the prenasalized and complex-onset series independently documented for Babouté (Vute) in the linguistic literature; the additional prenasalized labial-velar mgb, attested in some other datasets of the same ALCAM series, does not occur here.
A labialized consonant series is also attested — kw, gw, ngw (ŋgw), cw, jw, ndw — consistent with the phonemic labialization independently reported for Babouté (Vute) consonants. Vowel length is marked with a colon diacritic (:), consistent with the long/short vowel contrast reported for the language.
Three level tones are marked in the datasheet and mapping files:
High tone (H): á
Mid tone (M): ā
Low tone (L): à
Unmarked vowels represent tonally neutral or contextually determined syllables. Published descriptions of Babouté (Vute) additionally attest several contour-tone categories (e.g. mid-high rising, low-high, high-low, high-mid, high-low-high), which are not distinguished in this transcription; this simplification should be taken into account when using the dataset for tonal analysis.
The dataset was collected through a questionnaire designed to gather basic information about the Babouté lexicon and grammar. This was done as part of the Atlas Linguistique du Cameroun (ALCAM) project.
The dataset represents a linguistic questionnaire designed to elicit the basic lexicon and grammatical information of Babouté as spoken in Njoré.
The datasheet comprises 374 elicited lexical/example-sentence entries. The audio subfolder contains 245 audio files distributed across four recording batches (72, 43, 74 and 56 clips respectively; 5m 11s, 2m 37s, 5m 34s and 4m 08s of speech, measured directly from the audio files), for a total of 17 minutes 32 seconds of speech across 245 files. Unlike some other datasets in the same ALCAM series, no cross-batch redundancy (repeated sentences recorded across more than one batch) was found: all 245 recorded sentences are distinct, so the dataset contains 245 unique sentence-audio pairs.
The dataset comprises: 1) a datasheet (ALCAM_dataset_bàbútè-njorè.tsv) with lexical entries, French glosses, and example sentences with word-for-word parsing; 2) voice clips read by a native speaker of Babouté, organized into four recording batches under the audio subfolder; 3) sentence-to-audio mapping files (mapping.tsv, one per recording batch).
Datasheet (ALCAM_dataset_bàbútè-njorè.tsv):
OrigID: original number of lexical entry on paper questionnaire
EditID: modification of OrigID
FrenchRef: reference entry (originally provided in French)
FrenchComm: original comments about reference entry (FrenchRef)
French: lexical entry in French (overlaps with FrenchRef)
Note: note of researcher on the lexical entry
POS: part of speech
Class: noun class (where applicable); not populated in this release — Babouté (Vute) is a Bantoid language and does retain a reduced noun-class-related morphology in the broader literature, but this was not annotated as part of the present questionnaire
Morf: morphological attribute (e.g. plural, singular)
Var: dialect variant tag; not populated in this release, since the dataset represents a single, localized variety (see Variants above)
Word: lexical entry in Babouté (IPA)
CrossRef: cross-referencing of lexical entry number
FrenchEx: example sentence in French
LangEx: example sentence in Babouté (IPA)
LangExEdit: manual editing of LangEx
FrenchExEdit: edited French equivalent of FrenchEx
LangPars: word-for-word parsing in Babouté
LangParsEdit: editing of LangPars
FrenchPars: French equivalent of LangParsEdit
FrenchParsEdit: editing of FrenchPars
_ marks fields left empty in the source questionnaire. Of the 374 entries, 373 have a French gloss, 246 have a Babouté Word, 254 have a French example sentence, 191 have a Babouté example sentence, and 191 have word-for-word parsing. The audio subfolder separately provides recordings for 245 unique sentence-level items overall (see Size, above); a spot-check confirms 183 of these correspond exactly to LangEx entries in the datasheet (see Sample, below), but an exhaustive row-by-row correspondence between every audio clip and a specific datasheet entry has not been established.
The excerpt below shows 5 of the 374 datasheet entries (OrigID x1–x8) for which the Babouté example sentence (LangEx) has an exact, verified match in one of the audio-sentence mapping files, illustrating the parallel Babouté/French structure together with the aligned audio.
| OrigID | French | Word (IPA) | FrenchEx | LangEx (IPA) | Audio file |
|---|---|---|---|---|---|
| x1 | bouche | ɓō | elle a une petite bouche | ngá: ɓēé ɓō mə̀lénèé | 0f1dedbfed1a265cc7c4d96fbd1c74c7.mp3 |
| x2 | oeil | i: | ils se frottaient les yeux | ngáp tāá kpúktuì nw̄ íì:p | da2a54cc7f4ed69fddbfb1fed6f8c5c3.mp3 |
| x3 | tête | ngwé | il a une grosse tête et un long cou | ngé yāá ɓē ngwé jérē ɓēé mìur ngō̰ó̰ | b21953e6aeb6539271eec802cfec563b.mp3 |
| x4 | poil | nvùtié | ses poils sont noirs | nvuìtí ngé: ā síīp | 4bccb4ed668b0acc8778724df818919c.mp3 |
| x8 | oreille | tó | il se gratte derrière l'oreille | ngá: sòkmī njūúm tó: | 4781059ed1592eac279ef60b2bc794e6.mp3 |