Release Date: 8/24/2026
Format: WAV
Size: 83.42 MB
Share
This dataset contains Sheng audio and accurate human-transcript pairs within the mobile money domain. It covers a wide range of demographics with a rich distribution across gender/age group pairs. Audio was collected from individual respondents in-field (with phone microphone as the most common recording methodology). Speakers were asked to answer prompts/questions that focused specifically on mobile money usage, and provided free-form, unstructured responses. Rather than reading predetermined sentences, respondents spoke naturally about their experiences and usage. This methodology captures a diverse range of speakers, utterances, speaking styles, and recording environments. This makes the dataset ideal for the development of ASR modelling and applications where code-switching between Sheng, Swahili and English is expected, and speech is likely to occur in real-world environments where there is high acoustic variability. Currently only the 1 hour sample is available for download; we have a much larger corpus of data that will be published as a final dataset of 50 hours.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($500.00) platform fee paid to MDC.
Restrictions/Special Constraints
This dataset cannot be used to train any models or be used for applications that fit the following criteria: 1. Malicious 2. Unethical 3. Inappropriate Users must follow the data license as a governing usage guideline.
Forbidden Usage
You agree not to attempt to determine the identity of speakers in this dataset Any attempt to clone the voice or train models that imitate the speakers in this dataset is forbidden.
Intended Use
This dataset can be used for a broad set of applications, but is mainly intended for building robust ASR systems that support Sheng.
Sheng is an urban language variety widely spoken in Kenya, particularly among younger populations. It blends Swahili and English with vocabulary and influences from other Kenyan languages.
Sheng is characterized by:
Frequent code-switching
Rapidly evolving slang
Informal, everyday usage
The speakers selected to contribute to this dataset are from rural Kenya and self-qualified as Sheng speakers.
The current demographic distribution in the provided sample is:
| Demographic Category | Number of Unique Speakers |
|---|---|
| Men | 26 |
| Women | 23 |
| Ages 18ā24 | 24 |
| Ages 25ā34 | 24 |
| Undisclosed age group | 3 |
| Undisclosed gender group | 2 |
| Total Unique Speakers | 51 |
Note: Age and gender categories are independent demographic attributes, so their counts should not be summed together.
The dataset contains the following files and directories:
dataset/
āāā audio/
ā āāā *.wav
āāā metadata.csv
āāā README.md
audio/ ā Contains the full WAV audio recordings.
README.md ā Provides basic documentation for the dataset.
metadata.csv ā Contains the transcriptions and associated metadata for each recording.
In addition to the audio recordings and transcripts, the dataset includes metadata fields describing speaker characteristics and recording quality.
| Metadata Field | Description |
|---|---|
file_name | The audio file associated with the metadata row. Files can be located in the audio/ directory. |
transcription | The human-generated transcription of the audio recording. |
speaker_age_range | The age range of the speaker. |
speaker_gender | The gender of the speaker. |
audio_quality_labels | Human-assigned labels describing recording quality, including volume, background noise, audio consistency, and related characteristics. |
Audio was collected using mobile phone microphones through a structured data pipeline that requested WhatsApp voice notes from respondents. This approach was designed to facilitate participation among hard-to-reach populations.
Participants self-selected to take part in the collection process.
Participants were provided with a series of prompts and asked to record free-form spoken responses. Example prompts included:
Tell me about the last time you had a problem with your mobile service. What happened, and what did you do to try to fix it?
Tell me about the last time you contacted your mobile provider's customer service, such as Safaricom or Airtel. What was the problem, how did you contact them, and what happened?
After collection, submitted recordings were stored and reviewed for quality and submission accuracy.
The same individuals who produced the audio recordings were asked, through WhatsApp, to provide transcriptions of their own voice notes.
As a result, the dataset includes broad transcription coverage from speakers who self-declared as Sheng speakers.
To improve transcription accuracy and completeness, a cleaning pipeline was developed to perform several quality checks:
Semantic check Determines whether the user's transcription is coherent and meaningfully answers the original prompt.
Syntax check Reviews spelling, capitalization, grammar, and punctuation.
Audio alignment Uses a combination of human review and AI-generated transcription to check the submitted transcription for general alignment with the source audio and overall accuracy.
Following these checks, revisions were made where necessary to improve the completeness and quality of the final transcriptions.