License:
NOODL-1.0
Steward:
Institute of African Digital HumanitiesDataset ID:
cmrv275t1004lnu07vtchcapw
Release Date: 7/21/2026
Format: WAV, TSV
Size: 42.77 MB
Share
Badja-ALCAM-MultimodalDataset is a multimodal linguistic dataset dedicated to the documentation and technological enhancement of Badja (also spelled Badjia; dialect name Bakjo), one of the named dialects of Ewondo (also known as Beti, Yaunde; ISO 639-3: ewo), a Bantu language spoken primarily by the Beti people across Cameroon's Centre Region. Ewondo remains one of Cameroon's better-documented national languages overall, with an estimated 577,700 speakers (1982 figure, the most recent census-based estimate) and a role as a widely used trade language around Yaoundé. Despite this, its constituent dialects — including Badja — remain, like many locally defined varieties of Cameroon's national languages, sparsely represented in computational language resources at the level of the specific dialect rather than the language as a whole. The dataset comprises three closely aligned components: (i) a datasheet containing lexical entries and example sentences reflecting attested usage in Badja; (ii) audio recordings of a substantial subset of these entries, produced by a native speaker; and (iii) explicit audio-sentence mapping files enabling precise alignment between the textual and acoustic data. The dataset's primary added value lies in its explicit documentation of Badja (Badjia), a dialect of Ewondo that — despite Ewondo's comparatively large speaker base, its status as a major Beti-Fang trade language, and the existence of a harmonized practical orthography developed for national-language teaching — remains, at the level of this specific dialect, poorly represented in modern computational and pedagogical resources. The parallel availability of text in Badja and in French, together with aligned speech for a subset of entries, makes the dataset suitable for a range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), machine translation (MT), forced alignment and pronunciation modelling. From a methodological perspective, the dataset is designed to bridge the gap between language documentation and language technology. The datasheet's word-for-word parsing of both the Badja and French example sentences further supports morphological analysis and glossed-corpus studies. At the same time, the structured datasheet supports basic lexicographic and grammatical documentation, and pedagogical uses in teacher training and language revitalisation contexts. More broadly, the Badja-ALCAM-MultimodalDataset exemplifies an approach to Cameroonian language resources that documents linguistic variation at the level of specific dialects and speech communities, rather than only at the level of the named language (Ewondo) as a whole.
Licensing
Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)
https://licensingafricandatasets.com/nwulite-obodo-licenseRestrictions/Special Constraints
By downloading this dataset, you agree: - To use it for research and scientific use only - that you will not re-host or re-share this dataset
Forbidden Usage
You agree not to use the data for: determining the identity of the speaker(s) in the dataset; attempt to clone the voice or train models that imitate the speaker(s) in this dataset; Generative AI; reproduction; duplication; modification; augmentation; copying; distribution; transmission; display; sale; transfer; publication or creation of derivative works without the explicit permission of the legal owner of the dataset.
Intended Use
(a) Speech-related tasks: - Automatic speech recognition (ASR): Audio-text alignment allows the evaluation of speech recognition models for Badja. It should be noted that the sentences are transcribed using the IPA alphabet, rather than the harmonized practical orthography used for Ewondo in national-language teaching. - Text-to-speech (TTS): As the dataset contains clean sentence-audio pairs, it can also be used to evaluate speech synthesis or text-to-speech models. The use of IPA transcription rather than the standardized practical orthography should be taken into account when designing TTS experiments. - Speech-text alignment/forced alignment benchmarking: Fine-grained audio-text pairing provides ground truth for evaluating phoneme- or word-level aligners adapted to tonal Bantu languages of Cameroon. (b) Translation and multilingual tasks: - Machine translation (Badja ↔ French): The sentence-level alignment between Badja (`LangEx`) and French (`FrenchEx`) makes the datasheet usable as a small parallel corpus for evaluating translation models, with the caveat that the Badja side uses IPA transcription rather than the standardized orthography, and that aligned audio is only confirmed for a subset of entries. (c) Linguistic and lexicographic tasks: - Morphological analysis/glossed-corpus studies: The word-for-word parsing columns (`LangPars`, `FrenchPars`) support computational morphology, interlinear text (ILT) modelling, and grammar induction for Badja and other Beti-Fang varieties of Cameroon. - Lexicon documentation and part-of-speech tagging: The `Word`, `French` and `POS` columns are useful for building basic lexical resources and part-of-speech taggers for Badja, a dialect of Ewondo that remains largely absent from computational resources at the level of the specific dialect. - Language documentation and revitalisation: The aligned audio and text support pedagogical uses in teacher training and community-based language revitalisation efforts focused on the Badja speech community.
Badja (also spelled Badjia; dialect name Bakjo) is classified in standard typological work as a dialect of Ewondo (also known as Beti, Yaunde, Cameroonian; ISO 639-3: ewo; Guthrie code A.72), a Bantu language (Niger-Congo > Atlantic-Congo > Volta-Congo > Benue-Congo > Bantoid > Southern Bantoid > Bantu, Zone A > Beti) spoken by an estimated 577,700 people (1982 figure; more recent secondary estimates put the number of speakers above 500,000). Ewondo speakers are concentrated in the departments of Mfoundi, Mefou-et-Afamba, Mefou-et-Akono, Nyong-et-So'o and Nyong-et-Mfoumou in the Centre Region, and in part of the Océan department in the South Region. Ewondo is mutually intelligible, or closely related, to other Beti-Fang varieties such as Bulu, Eton and Fang, and is widely used as a trade language in and around Yaoundé. This dataset documents the Badja (Badjia) dialect specifically.
Published descriptions of Ewondo list a substantial number of named dialects alongside Badjia (Bakjo) — including Bafeuk, Bemvele (Mvele, Yezum, Yesoum), Bane, Beti, Enoah, Evouzom, Mbida-Bani, Mvete, Mvog-Niengue, Omvang, Yabekolo (Yebekolo), Yabeka and Yabekanga — reflecting the fine-grained, clan- and lineage-based dialectal fragmentation characteristic of Ewondo/Beti. This release documents a single, localized dialect (Badja/Badjia) as elicited from a native speaker; no further internal sub-variety split is documented within the dataset, and the precise village- or locality-level provenance of the variety documented here has not been recorded and should be clarified with the dataset's creator before downstream use.
The writing system used for the transcription of Badja in this dataset is the International Phonetic Alphabet (IPA), as reflected in the Word, LangEx and LangPars columns of the datasheet and in the sentence field of the audio-sentence mapping files. The phonological inventory below draws on the 343 datasheet entries and the 326 audio-aligned sentences (the mapping.tsv files across the dataset's four recording batches). A harmonized practical orthography for Ewondo, developed for national-language teaching and Bible translation, is well established, but is not the transcription system used in this dataset.
The data attests eight oral vowel qualities: a, e, ɛ, ə, i, o, ɔ, u, matching the vowel inventory reported for Ewondo in the wider literature. Vowel length is occasionally marked with a colon diacritic (5 tokens); no nasalized vowels are systematically marked in the transcribed data. A combining tilde-below diacritic (◌̰), which in IPA conventionally marks creaky voice, occurs on 73 tokens across the Word, LangEx and LangPars columns — the same convention observed in other datasets elicited with the same ALCAM questionnaire; its precise phonetic conditioning has not been established here and should be clarified with the dataset's creator before downstream use.
The core simple-consonant inventory attested in the datasheet and mapping files is: b, d, g, j, k, l, m, n, p, r, s, t, v, w, y, z. Notably, f and h, both attested in published Ewondo consonant inventories, do not occur in the transcribed data.
Two affricates are attested: dʒ (voiced postalveolar, 4 tokens, occurring only in prenasalized position, i.e. always as ndʒ) and dz (voiced alveolar, 10 tokens, occasionally prenasalized as ndz, 3 tokens).
Prenasalized consonants are well attested: mb (75 tokens), nd (31), ŋg (88) and nz (76), consistent with the prenasalized consonant series widely reported for Ewondo and other Beti-Fang languages.
A labial-velar stop with a tie bar, k͡p, is attested (18 tokens); a corresponding tied voiced labial-velar g͡b does not occur — the plain sequence "gb" is found only twice and more likely reflects an adjacent, non-clustered g and b than a phonemic labial-velar.
Labialized consonants are only sparsely attested (sw, 3 tokens; tw, 2 tokens) and do not form a fully developed phonemic series comparable to the labialized series documented for some other languages in the same ALCAM series.
Word-final voiceless stops (t, p, k) frequently carry a superscript ʳ marking an audible release (117 tokens, e.g. títʳ "animals", kápʳ). A small number of forms instead mark aspiration (ʰ, 2 tokens, e.g. kʰu, pʰu) or an explicitly unreleased stop (◌̚, 2 tokens, e.g. bòk̚), suggesting variability in the realization of word-final stops that has not been further analysed here.
Two further symbols recur without a clearly documented function and should be clarified with the dataset's creator before downstream use: a diaeresis on w or n (6 tokens, e.g. ásẅɔ, ètün), and a word-final q restricted to three lexical entries (6 tokens total, e.g. mə́ndʒɔ̀q, bə̀ndʒɔ̀qʳ). An apostrophe (') also recurs on 5 tokens; as with the symbols above, its precise grammatical or phonetic function has not been formally documented.
Three level tones are marked in the datasheet and mapping files:
High tone (H): á
Mid tone (M): ā
Low tone (L): à
Unmarked vowels represent tonally neutral or contextually determined syllables. A rising tone, marked with a caron (ǎ), occurs on only 2 tokens; no falling (circumflex-marked) contour tone is attested in this data, even though the harmonized orthographic conventions used for Ewondo in national-language teaching provide for both rising and falling tone marks. This simplification, and the very limited attestation of contour tone generally, should be taken into account when using the dataset for tonal analysis.
The dataset was collected through a questionnaire designed to gather basic information about the Badja lexicon and grammar. This was done as part of the Atlas Linguistique du Cameroun (ALCAM) project.
The dataset represents a linguistic questionnaire designed to elicit the basic lexicon and grammatical information of Badja (Badjia).
The datasheet comprises 343 elicited lexical/example-sentence entries. The audio subfolder contains 326 audio files distributed across four recording batches (91, 99, 40 and 96 clips respectively; 4m 56s, 3m 50s, 2m 04s and 5m 26s of speech, measured directly from the audio files), for a total of 16 minutes 18 seconds of speech across 326 files. No cross-batch redundancy (repeated sentences recorded across more than one batch) was found: all 326 recorded sentences and all 326 audio filenames are distinct, so the dataset contains 326 unique sentence-audio pairs.
The dataset comprises: 1) a datasheet (AlCAM_dataset_badjia.tsv) with lexical entries, French glosses, and example sentences with word-for-word parsing; 2) voice clips read by a native speaker of Badja, organized into four recording batches under the audio subfolder; 3) sentence-to-audio mapping files (mapping.tsv, one per recording batch).
Datasheet (AlCAM_dataset_badjia.tsv):
OrigID: original number of lexical entry on paper questionnaire
EditID: modification of OrigID
FrenchRef: reference entry (originally provided in French)
FrenchComm: original comments about reference entry (FrenchRef)
French: lexical entry in French (overlaps with FrenchRef)
Note: note of researcher on the lexical entry
POS: part of speech
Class: noun class (where applicable); not populated in this release — Ewondo is a Bantu language with a productive noun-class system in the broader literature, but this was not annotated as part of the present questionnaire
Morf: morphological attribute (e.g. plural, singular)
Var: dialect variant tag; not populated in this release, since the dataset represents a single, localized dialect (see Variants above)
Word: lexical entry in Badja (IPA)
CrossRef: cross-referencing of lexical entry number
FrenchEx: example sentence in French
LangEx: example sentence in Badja (IPA)
LangExEdit: manual editing of LangEx
FrenchExEdit: edited French equivalent of FrenchEx
LangPars: word-for-word parsing in Badja
LangParsEdit: editing of LangPars
FrenchPars: French equivalent of LangParsEdit
FrenchParsEdit: editing of FrenchPars
_ marks fields left empty in the source questionnaire. Of the 343 entries, 342 have a French gloss, 334 have a Badja Word, 263 have a French example sentence, 259 have a Badja example sentence, and 259 have word-for-word parsing. The audio subfolder separately provides recordings for 326 unique sentence-level items overall (see Size, above); a spot-check (allowing for Unicode normalization of combining diacritics) confirms 7 of these correspond exactly to LangEx entries in the datasheet (see Sample, below), but an exhaustive row-by-row correspondence between every audio clip and a specific datasheet entry has not been established.
The excerpt below shows 5 of the 343 datasheet entries (OrigID x77, x183, x203, x204, x205) for which the Badja example sentence (LangEx) has an exact, verified match (after Unicode normalization) in one of the audio-sentence mapping files, illustrating the parallel Badja/French structure together with the aligned audio.
| OrigID | French | FrenchEx | LangEx (IPA) | Audio file |
|---|---|---|---|---|
| x77 | viande | c'est de la bonne viande | m̀bɛ̀mbɛ̀ tít | e9e63376beb79033ecb94c39759d8bcc.wav |
| x183 | dire | je n'entends pas ce que tu dis | mɔ̀ tá ánɔ̀ŋə yɔ́mə wàáyù | 6df84a68ac1333ab4220fe6fc8f81cf6.wav |
| x203 | de | elle vient du marigot | àátɛ̀m bàŋ ájí | 83ff2a87f86fbb9ba344a9bf2f677f7d.wav |
| x204 | dans | il y a de l'eau dans le marigot | mə̀jí mə́n(ə) ā́jī | a99f90a49c164cf0e9a643c0a8fa810b.wav |
| x205 | derrière | il est assis derrière la case | ànə̀ ǹláŋ ámbús pāɣā | 31ba965ece0d67bc13e59bb8f1201ac0.wav |