Release Date: 7/21/2026
Format: MP3, TSV
Size: 235.91 MB
Share
This dataset is a subset of the larger Arabic Mozilla Common Voice Scripted Speech 26.0 dataset. It contains only validated audio (at least one upvote and 0 downvotes) from users whose self-identified gender is Male/Masculine.
Restrictions/Special Constraints
NA
Forbidden Usage
You agree not to attempt to determine the identity of speakers in this dataset.
Intended Use
Speech technology; Linguistics research
This dataset was created by filtering the train, dev, and test files from the v26 Arabic Mozilla Common Voice Scripted Speech dataset with the following conditions:
the value of the up_votes field is a number greater than 0.
the value of the down_votes field is 0.
the value of the gender column is "Male_Masculine"
The resulting dataset includes 2,720 clips in the training set, 2,996 clips in the dev set, and 3,560 clips in the test set, totaling approximately 11 hours, 7 minutes of audio.
The dataset follows the Mozilla Common Voice format: The clips directory contains all of the .mp3 files, and there is a separate tsv file for each data partition, containing the following fields:
client_id
path
sentence_id
sentence
sentence_domain
up_votes
down_votes
age
gender
accents
variant
localesegment
| Category | Train (n) | Train (%) | Dev (n) | Dev (%) | Test (n) | Test (%) |
|---|---|---|---|---|---|---|
| Not Specified | 0 | 0% | 77 | 3% | 2 | 0% |
| Teens | 0 | 0% | 171 | 6% | 175 | 5% |
| Twenties | 1,654 | 61% | 1,077 | 36% | 1,908 | 54% |
| Thirties | 1,066 | 39% | 1,384 | 46% | 1,141 | 32% |
| Fourties | 0 | 0% | 287 | 10% | 268 | 8% |
| Fifties | 0 | 0% | 0 | 0% | 19 | 1% |
| Sixties | 0 | 0% | 0 | 0% | 44 | 1% |
| Nineties | 0 | 0% | 0 | 0% | 3 | 0% |
| Total | 2,720 | 100% | 2,996 | 100% | 3,560 | 100% |