License:
CC-BY-NC-4.0
Steward:
CommunityDataset ID:
cmuo1rfsj01runn073id66vg7
Release Date: 9/30/2026
Format: WAV, TSV
Size: 3.35 GB
The Central Kanuri Speech Dataset is a community-contributed speech corpus for the Central Kanuri (kau) language of the Lake Chad Basin, collected through the TWB Voice platform, with CIATECH Africa leading implementation and community engagement. The dataset contains 3,576 recordings totalling approximately 12 hours of approved audio, contributed by 24 speakers across two collection workflows: 2,134 read-speech clips, where the text read aloud is the prompt, and 1,442 freeform spoken answers to spoken questions. Every clip carries the platform's recording-approval status, with reviewed recordings published as approved, rejected, or pending rather than withholding audio. 2,982 recordings have an associated transcript, while 595 freeform recordings were not transcribed and are published with their prompt, duration, review outcome, and speaker metadata only. Audio was captured on contributors' devices and is included in full, comprising 3,151 files at 48 kHz and 425 files at 44.1 kHz. The dataset is intended to support automatic speech recognition (ASR) research, prototyping, and evaluation for an underserved language in the Lake Chad Basin. It is unlikely to be sufficient on its own for training an end-to-end ASR system.
Licensing
Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)
https://spdx.org/licenses/CC-BY-NC-4.0.htmlRestrictions/Special Constraints
By downloading the data, the user becomes an independent data controller and is responsible for complying with applicable data protection laws (including GDPR), responding to data subject rights requests, and reporting any data breaches. Users must implement appropriate technical and organizational security measures and retain data only as long as necessary for the intended purpose. Voice recordings are inherently biometric: users must not attempt to re-identify contributors, must not use the data for speaker identification or biometric profiling, and must not combine it with other data to identify individuals.
Forbidden Usage
Attempting to determine or confirm the identity of any contributor. Re-hosting or re-sharing the dataset. Any commercial use.
| Field | Description |
|---|---|
| Dataset name | Central Kanuri Speech Dataset |
| Language | Central Kanuri |
| ISO 639-3 | kau |
| Region | Lake Chad Basin |
| Collection platform | TWB Voice |
| Implementation lead | CIATECH Africa |
| Contributors | 24 |
| Total recordings | 3,576 |
| Total approved audio | 12 hours |
| Transcribed recordings | 2,982 |
| Untranscribed freeform recordings | 595 |
| Primary use | Automatic Speech Recognition (ASR) |
| Dataset type | Community-contributed speech corpus |
The dataset contains 3,576 recordings collected through two workflows:
The two workflows provide complementary speech data: read speech provides controlled linguistic content, while freeform speech captures more natural conversational responses.
Audio was recorded on contributors' personal devices and is included in full.
| Audio characteristic | Number of files |
|---|---|
| 48 kHz | 3,151 |
| 44.1 kHz | 425 |
| Total | 3,576 |
The dataset contains approximately 12 hours of approved audio.
A total of 2,982 recordings carry a transcript.
The remaining 595 freeform recordings were not transcribed. These records are nevertheless published with available metadata, including:
For read-speech recordings, the text read aloud corresponds to the collection prompt.
Each recording carries the recording-approval status generated through the TWB Voice collection workflow.
Reviewed recordings are represented using the platform's outcomes:
The dataset retains the audio and associated metadata rather than withholding recordings solely on the basis of review status.
The dataset is intended to support:
The dataset should be considered a community-contributed low-resource speech corpus.
Although it provides useful material for Central Kanuri ASR research, it is unlikely to be sufficient on its own to train a robust end-to-end ASR system.
Researchers should consider combining it with additional appropriately licensed Central Kanuri speech resources where possible.
The dataset also contains both read and freeform speech, meaning that the linguistic characteristics and transcription availability vary across recordings.
Each recording is associated with metadata describing the collection and speaker where available.
Typical metadata includes:
Users should account for differences in:
For ASR experiments, researchers are encouraged to create explicit training, development, and test partitions and to ensure that speakers do not overlap across evaluation splits.
Recommended evaluation metrics include:
For Central Kanuri, character-level evaluation may be particularly useful where word segmentation and orthographic conventions create challenges for word-level comparison.
The dataset contributes a community-generated speech resource for Central Kanuri, an underserved language of the Lake Chad Basin.
It provides a foundation for further data collection, transcription, ASR experimentation, and development of locally relevant speech technologies.
Researchers and developers should respect the privacy, consent, attribution, and data-use conditions associated with the collection.
The dataset should be used primarily for research, evaluation, and responsible development of language technologies that benefit Central Kanuri-speaking communities.
This is the canonical release: Central Kanuri Speech Dataset on Hugging Face, which is also the source of this package.