License:
NOODL-1.0
Steward:
Institute of African Digital HumanitiesDataset ID:
cmrw1urb2001nnu07i6vv0vea
Release Date: 7/22/2026
Format: WAV, TSV
Size: 976.53 MB
Share
This dataset comprises audio recordings of Bamun (Shupamem) speech aligned with textual transcriptions. It is a female-voice companion to the previously published Bamun-TTS-Dataset (male voice), and is structured into 38 folders, each containing audio files and a corresponding audio-text mapping file. The audio clips are short, typically ranging from 1 to 10 seconds (with a small number of clips up to about 30 seconds), and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file. The textual content used in this dataset originates from transcriptions of oral narratives documenting personal histories related to German colonisation in Cameroon. These texts were segmented into short utterances suitable for read speech and TTS modelling. The same textual material was used for the companion male-voice dataset.
Licensing
Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)
Restrictions/Special Constraints
- For research and scientific use only - You agree not to re-host or redistribute this dataset
Forbidden Usage
You agree not to use the data for: - Generative AI - Voice cloning or speaker imitation - Reproduction, duplication, modification, or redistribution - Commercial use without explicit permission
Intended Use
This dataset is intended for the training and evaluation of Text-to-Speech (TTS) systems for the Bamun language. It aims to support: - Language revitalisation - Development of speech technologies for under-served African languages - Educational applications in multilingual contexts - Research on reading-skill acquisition and usability of AGLC-based orthography
Bamun or Shüpamom/Shupamem is a Bantu-Grassfield language spoken in the Noun Division, West Region in Cameroon.
The Bamun language is quite homogeneous within their indigenous territory, the Noun Administrative Division. However, the Administrative Atlas of Cameroon's Languages (Breton and Bikia Fohtung, 1991) indicates a few "islands" outside the Noun Department where the Bamun language exhibits minor variations. These include Bapi in the Mifi Division in the West Region and Bamalang and Bangolan in the Mezam Division in the Northwest Region.
The vowel inventory reflected in the dataset is: i, e, ɛ, a, ɔ, o, u, ʉ, ə
The vowel ə / ә is particularly frequent and functions as a central vowel.
The consonant system includes the following simple consonants: b, d, f, g, h, j, k, l, m, n, ŋ, p, r, s, t, v, w, y, z
Complex and cluster-like consonants attested include: mb, nd, nk, ng, nt, nj, mf, kp
Digraphs: sh, gh
The transcription encodes lexical tone using diacritics, corresponding to standard tonal categories:
High tone (H): marked with acute accent (á, é, ɔ́, ʉ́, ŋ́)
Low tone (L): marked with grave accent (à, è, ɔ̀)
Mid tone (M): marked with macron (ā, ē)
Rising tone (LH): marked with caron (ǎ, ě, ɔ̌)
Falling tone (HL): marked with circumflex (â, ê)
This dataset originates from audio recordings documenting personal histories of German colonisation. These recordings were made in the early eighties as part of a research project led by Prince (Professor) Koum A Ndoumbe III.
Abdou Salam Ntieche Fifen created the transcriptions associated with this dataset. The transcriptions were made in 2017 under the coordination of the AfricAvenir Foundation.
For the purpose of creating this dataset, the textual material was segmented into short utterances and aligned with corresponding audio recordings to support TTS modelling. This female-voice recording uses the same segmented textual material as the companion Bamun-TTS-Dataset (male voice).
This dataset is derived from prompted speech in the form of directed interviews. The content reflects personal narratives related to colonial history in Cameroon.
The dataset has been transformed into read-style segmented speech suitable for speech synthesis tasks.
This dataset was recorded by a female native Bamun speaker, reading the same textual material as the male speaker of the companion Bamun-TTS-Dataset (male voice).
During recording, the female speaker appeared less proficient at reading the AGLC-based orthography (Alphabet Général des Langues Camerounaises / General Alphabet of Cameroon's Languages) than the male speaker of the companion dataset. This discrepancy in reading fluency was accidental and was not orchestrated or planned. The decision was made to proceed with and publish the recording as captured, since it may carry independent research value beyond TTS: AGLC is a Latin-based script incorporating IPA characters and tone diacritics, and the development of AGLC-based orthographic standards remains a recurring subject of debate among Cameroonian language communities. Data such as this may support research into reading-skill acquisition and orthography usability for AGLC-based writing systems.
Users comparing or combining this dataset with the male-voice dataset should be aware of this difference in speaker reading fluency.
The dataset is composed of 38 folders containing audio clips and corresponding mapping files.
Each folder contains between 95 and 100 audio files. Individual audio clips typically range from 1 to 10 seconds, with a small number of clips (about 6% of the dataset) extending up to roughly 30 seconds.
Folder-level durations range from approximately 5 minutes to 21 minutes of audio. In total, the dataset contains 3,718 audio clips amounting to approximately 5 hours 4 minutes of segmented Bamun speech data.
A detailed breakdown of durations and file counts per folder is provided in the accompanying duration report.
Each folder in the dataset contains:
A collection of audio files in WAV format
A tab-separated mapping file linking each audio file to its transcription
Each line in the mapping file follows the format:
audio_filename.wav transcription
The dataset is designed for TTS pipelines requiring paired audio-text data.
4338880fab47658c4352d8b988bd7c3e.wav | Mí u tóóshә́ ŋwәt ru, nzíé yúá yʉ́ә́ u púá' yírә́ nә́.
3857b094a7e57be4fb26c9ad305f4d07.wav | Í nzie Li shá?
5778209f616657c7a6ca064f8bfabbad.wav | Li shú, nә nguu yúá u púá' yírә́ nә́, nә́ nguu yúá u yí nә́.
ebc42edb24b6f8ece9f528646826b0f0.wav | Ndǔ lʉ́m mú, ká u lá' nә́, tә njʉ́ nә́.
8519444c581caeb3bd2a5195a63f36b8.wav | Pә́ ka pә́ nkʉ́ ndɛ́t saangǎm