Release Date: 7/21/2026
Format: MP3, TSV
Size: 185.04 MB
Share
This dataset is a subset of the larger Arabic Mozilla Common Voice Scripted Speech 26.0 dataset. It contains only validated audio (at least one upvote and 0 downvotes) from users who's self-identified gender is Female/Feminine.
Restrictions/Special Constraints
NA
Forbidden Usage
You agree not to attempt to determine the identity of speakers in this dataset.
Intended Use
Speech technology; Linguistics research
This dataset was created by filtering the train, dev, and test files from the v26 Arabic Mozilla Common Voice Scripted Speech dataset with the following conditions:
the value of the up_votes field is a number greater than 0.
the value of the down_votes field is 0.
the value of the gender column is "Female_Feminine"
The resulting dataset includes 3,620 clips in the training set, 3,055 clips in the dev set, and 1,356 clips in the test set, totaling approximately 9 hours, 22 minutes of audio.
The dataset follows the Mozilla Common Voice format: The clips directory contains all of the .mp3 files, and there is a separate tsv file for each data partition, containing the following fields:
client_id
path
sentence_id
sentence
sentence_domain
up_votes
down_votes
age
gender
accents
variant
localesegment
| Category | Train (n) | Train (%) | Dev (n) | Dev (%) | Test (n) | Test (%) |
|---|---|---|---|---|---|---|
| Not Specified | 0 | 0% | 413 | 14% | 23 | 2% |
| Teens | 811 | 22% | 126 | 4% | 185 | 14% |
| Twenties | 2,809 | 78% | 2,127 | 70% | 1,040 | 77% |
| Thirties | 0 | 0% | 389 | 13% | 105 | 8% |
| Fourties | 0 | 0% | 0 | 0% | 2 | 0% |
| Sixties | 0 | 0% | 0 | 0% | 1 | 0% |
| Total | 3,620 | 100% | 3,055 | 100% | 1,356 | 100% |