Release Date: 7/21/2026
Format: MP3, TSV
Size: 121.80 MB
Share
This dataset is a subset of the larger Dutch Mozilla Common Voice Scripted Speech 26.0 dataset. It contains only validated audio (at least one upvote and 0 downvotes) from users whose self-identified accent corresponds to Netherlands Dutch and whose self-identified gender is Female/Feminine.
Restrictions/Special Constraints
NA
Forbidden Usage
You agree not to attempt to determine the identity of speakers in this dataset.
Intended Use
Speech technology; Linguistics research
This dataset was created by filtering the train, dev, and test files from the v26 Dutch Mozilla Common Voice Scripted Speech dataset with the following conditions:
the value of the up_votes field is a number greater than 0.
the value of the down_votes field is 0.
the value of the gender column is "female_feminine".
the value of the accents column is one of the following:
{Nederlands met Texels accent,
Nederlands Nederlands,
Nederlands westfries}
The resulting dataset includes 2,714 clips in the training set, 1,055 clips in the dev set, and 556 clips in the test set, totaling approximately 5 hours, 53 minutes of audio.
The dataset follows the Mozilla Common Voice format: The clips directory contains all of the .mp3 files, and there is a separate tsv file for each data partition, containing the following fields:
client_id
path
sentence_id
sentence
sentence_domain
up_votes
down_votes
age
gender
accents
variant
localesegment
| Category | Train (n) | Train (%) | Dev (n) | Dev (%) | Test (n) | Test (%) |
|---|---|---|---|---|---|---|
| Not Specified | 0 | 0% | 0 | 0% | 5 | 1% |
| Teens | 0 | 0% | 13 | 1% | 18 | 3% |
| Twenties | 1 | 0% | 119 | 11% | 349 | 63% |
| Thirties | 105 | 4% | 8 | 1% | 45 | 8% |
| Fourties | 1,674 | 62% | 452 | 43% | 54 | 10% |
| Fifties | 449 | 17% | 451 | 43% | 57 | 10% |
| Sixties | 485 | 18% | 12 | 1% | 28 | 5% |
| Total | 2,714 | 100% | 1,055 | 100% | 556 | 100% |