Release Date: 7/20/2026
Format: TSV, MP3
Size: 330.28 MB
Share
This dataset is a derived, accent-filtered subset of the Mozilla Common Voice Scripted Speech 26.0 English release (`cv-corpus-26.0-2026-06-12`, locale `en`). It keeps only validated rows (present in the source `validated.tsv`) whose self-declared `accents` field is *exactly* `"Irish English"`. Common Voice carries a separate, distinct predefined tag `"Northern Irish"` (part of the UK); this subset is deliberately restricted to the Republic-of-Ireland `Irish English` tag only. No gender or age filter is applied. The result is split speaker-disjoint into train/dev/test (no `client_id` appears in more than one split).
Restrictions/Special Constraints
N/A
Forbidden Usage
- You agree not to attempt to determine the identity of speakers. - You agree not to rehost this dataset.
Ethical Review
Source recordings and self-declared demographics (age, gender, accent) were volunteered by Common Voice contributors under Common Voice's own consent and contribution terms. This package only filters and repackages a subset; it does not collect new data or alter consent terms.
Intended Use
Research and development of English ASR systems, with emphasis on a clean, single-tag Irish English (Republic of Ireland) accent slice.
Splits and speakers: 9,528 validated clips total — train 9,031 clips / 193 speakers, dev 255 clips / 20 speakers, test 242 clips / 7 speakers. Split is speaker-disjoint (no client_id appears in more than one split), proportioned ~94.78/2.68/2.54 by clip count via a greedy per-speaker balancing algorithm targeting 95/2.5/2.5 (speakers have uneven clip counts, so exact ratios aren't achievable while keeping speakers disjoint).
Clip duration: 13.01 hours total (46,850 seconds) across 9,528 clips. Per-split: train 12.36h, dev 0.38h, test 0.27h. Per-clip: mean 4.92s, median 4.87s, min 0.82s, max 13.94s, std 1.75s.
Gender distribution: male_masculine 6,159; female_feminine 2,772; blank/unspecified 597.
Age distribution (shipped subset, self-declared): thirties 3,015; twenties 2,162; fifties 1,859; fourties 1,558; sixties 374; teens 294; blank/unspecified 235; seventies 31.
Closed-form accent matching (QA): the accents field is pipe-delimited multi-select in the source corpus, but this subset was deliberately filtered to the closed-form (exact-match) rule — 100% of shipped clips have accents == exactly "Irish English", no co-occurring tags.
Open-form vs. closed-form comparison (at execution time): open-form (accents contains "Irish English") = 9,725 clips / 238 speakers; closed-form (accents == exactly "Irish English") = 9,528 clips / 220 speakers / 13.01h. Closed-form drops 197 clips (2.03%).
Disambiguation from Northern Irish: Common Voice carries two distinct predefined accent tags in this region — Irish English (Republic of Ireland, 9,528 clips closed-form) and Northern Irish (Northern Ireland, part of the UK, 10,546 clips closed-form). These are not variants of one accent. This subset's locale is en-IE with only Irish English in scope. Northern Irish population is already represented in the British English release.
Full-corpus accent x gender and age x gender cross-tabs are in stats.md.
Derivation methodology: derived from the Mozilla Common Voice Scripted Speech 26.0 English release (cv-corpus-26.0-2026-06-12, locale en). Filter: validated rows only (present in the source validated.tsv), whose self-declared accents field is exactly "Irish English" (closed-form/exact match, no co-occurring tags, distinct from the separate "Northern Irish" tag). No gender or age filter applied. Speaker-disjoint train/dev/test split via a greedy per-speaker assignment balancing toward 95/2.5/2.5 targets. Only the referenced MP3 clips were extracted from the ~2.58M-clip, ~89GB source archive (not the full archive) via a single sequential streaming pass over the source tar.gz.