Release Date: 7/24/2026
Format: TSV, MP3
Size: 32.34 MB
Share
This dataset is a derived, accent-filtered subset of the Mozilla Common Voice Scripted Speech 26.0 English release (`cv-corpus-26.0-2026-06-12`, locale `en`). It keeps only validated rows (present in the source `validated.tsv`) whose self-declared `accents` field exactly matches one of 21 curated allowlist values identifying Southern American English self-identification (closed-form/exact match against a hand-reviewed allowlist, not a keyword/substring match). No gender or age filter is applied. The result is split speaker-disjoint into train/dev/test (no `client_id` appears in more than one split).
Restrictions/Special Constraints
- You agree not to attempt to determine the identity of speakers. - You agree not to rehost this dataset. - You agree not to neither create nor publish voice-cloning systems with this data
Forbidden Usage
- You agree not to attempt to determine the identity of speakers. - You agree not to rehost this dataset. - You agree not to neither create nor publish voice-cloning systems with this data
Ethical Review
Source recordings and self-declared demographics (age, gender, accent) were volunteered by Common Voice contributors under Common Voice's own consent and contribution terms. This package only filters and repackages a subset; it does not collect new data or alter consent terms.
Intended Use
Research and development of English ASR systems, with emphasis on a Southern American English accent slice.
Closed-form allowlist (21 values): United States English|southern United States; United States English|southern United States|New Orleans dialect; United States English|South Texas|Slightly effeminate|Conversational; United States English|Texas; United States English|United States Southeastern English (Charlotte NC); United States English|southern draw; United States English|thicc southern drawl; United States English|Southern United States English|Lightly Southern; United States English|i have some pronunciation issues because of oral surgery and a hidden southern accent; United States English|East Texas|Texan; United States English|Kentucky united states accent|southern usa; United States English|Southern Texas Accent; United States English|Midwest USA speech blended with South Texas USA speech; United States English|Southern Appalachian English|Southern United States English; United States English|southern, formal, sultry; United States English|Deep South US; United States English|kentucky; United States English|Southern with a stutter; United States English|Southern United States English; United States English|East Texas; United States English|Southern|eastern us|Appalachia.
Splits and speakers: 1,175 validated clips total — train 709 clips / 23 speakers, dev 374 clips / 1 speaker, test 92 clips / 1 speaker. Split is speaker-disjoint (no client_id appears in more than one split), proportioned 60.34/31.83/7.83 by clip count via a greedy per-speaker balancing algorithm targeting 95/2.5/2.5. This led to dev and test each necessarily consist of exactly one speaker's clips.
Clip duration: 1.79 hours total (6,439 seconds) across 1,175 clips. Per-split: train 1.15h, dev 0.49h, test 0.15h. Per-clip: mean 5.48s, median 5.36s, min 1.94s, max 10.55s, std 1.52s.
Gender distribution: male_masculine 570; female_feminine 464; blank/unspecified 141.
Age distribution (shipped subset, self-declared): twenties 477; fifties 475; blank/unspecified 131; sixties 82; fourties 6; thirties 4.
Closed-form accent matching (QA): the accents field is pipe-delimited multi-select in the source corpus, but this subset was deliberately filtered to the closed-form (exact-match) rule against the 21-value allowlist — 100.0000% (1,175/1,175) of shipped clips have accents matching exactly one allowlist value, no drift, no partial matches.
Accent value breakdown: unlike the multi-select open-form subsets in this project, every shipped clip's accents field matches exactly one of the 21 allowlist values (no exploded/presence-based double counting). 20 of the 21 allowlist values have at least one shipped clip:
| Accent value | Clips | % of clips | Duration (h) |
|---|---|---|---|
United States English|southern United States | 551 | 46.89% | 0.88 |
United States English|southern United States|New Orleans dialect | 374 | 31.83% | 0.49 |
United States English|South Texas|Slightly effeminate|Conversational | 92 | 7.83% | 0.15 |
United States English|Texas | 83 | 7.06% | 0.15 |
United States English|thicc southern drawl | 9 | 0.77% | 0.01 |
United States English|southern draw | 9 | 0.77% | 0.02 |
United States English|United States Southeastern English (Charlotte NC) | 9 | 0.77% | 0.01 |
United States English|Southern United States English|Lightly Southern | 7 | 0.60% | 0.01 |
United States English|East Texas|Texan | 5 | 0.43% | 0.01 |
United States English|i have some pronunciation issues because of oral surgery and a hidden southern accent | 5 | 0.43% | 0.01 |
United States English|Kentucky united states accent|southern usa | 4 | 0.34% | 0.01 |
United States English|Midwest USA speech blended with South Texas USA speech | 4 | 0.34% | 0.01 |
United States English|Southern Appalachian English|Southern United States English | 4 | 0.34% | 0.01 |
United States English|Southern Texas Accent | 4 | 0.34% | 0.00 |
United States English|southern, formal, sultry | 4 | 0.34% | 0.01 |
United States English|Deep South US | 3 | 0.26% | 0.01 |
United States English|kentucky | 3 | 0.26% | 0.01 |
United States English|Southern with a stutter | 3 | 0.26% | 0.01 |
United States English|East Texas | 1 | 0.09% | 0.00 |
United States English|Southern United States English | 1 | 0.09% | 0.00 |
Full-corpus accent x gender and age x gender cross-tabs are in stats.md.
Derivation methodology: derived from the Mozilla Common Voice Scripted Speech 26.0 English release (cv-corpus-26.0-2026-06-12, locale en). Filter: validated rows only (present in the source validated.tsv), whose self-declared accents field exactly matches one of the 21 curated allowlist values (closed-form/exact match, no keyword/substring matching). No gender or age filter applied. Speaker-disjoint train/dev/test split via a greedy per-speaker assignment balancing toward 95/2.5/2.5 targets, largest speakers placed first. Only the referenced MP3 clips were extracted from the ~2.58M-clip, ~89GB source archive (not the full archive) via a single sequential streaming pass over the source tar.gz.
Small-corpus and speaker-concentration note: at 1,175 clips / 25 speakers / 1.79 hours, this is the smallest of the ten Common Voice accent-subset MDC submissions in this project, and severely speaker-concentrated (top 2 of 25 speakers = 78.0% of clips). Dev and test splits each consist of exactly one speaker's clips as a direct consequence.