Release Date: 9/17/2026
Format: MP3
Size: 225.34 MB
A collection of spontaneous responses to questions in Gorani (Gorani).
Restrictions/Special Constraints
None provided.
Forbidden Usage
It is forbidden to attempt to determine the identity of speakers in the Common Voice datasets. It is forbidden to re-host or re-share this dataset.
Intended Use
This dataset is intended to be used for training and evaluating automatic speech recognition (ASR) models. It may also be used for applications relating to computer-aided language learning (CALL) and language or heritage revitalisation.
hac)This datasheet is for sps-corpus-5.0-2026-09-11 of the Mozilla Common Voice Spontaneous Speech dataset for Gorani [Gorani - hac]. The dataset contains 421 clips representing 11.19 hours of recorded speech (0.37 hours validated) from 19 speakers.
The dataset clips are categorised by transcription status and training-set assignment. The following tables summarise the distribution.
| Bucket | Clips | % |
|---|---|---|
| Transcribed & Validated | 14 | 3.3% |
| Transcribed & Pending | 13 | 3.1% |
| Not transcribed | 394 | 93.6% |
| Bucket | Clips | % |
|---|---|---|
| Train | 0 | 0.0% |
| Dev | 0 | 0.0% |
| Test | 0 | 0.0% |
| Unassigned | 421 | 100.0% |
Training split coverage: 0 of 14 transcribed & validated clips (0.0%)
| Bucket | Clips | % |
|---|---|---|
| Validated | 14 | 51.9% |
| Pending | 13 | 48.1% |
| Edited | 0 | 0.0% |
There follows a randomly selected sample of questions used in the corpus.
نامۆ هیچ هەسارێڤ یا سوەرێ فەلەکیێڤی مزانی؟
لا شمە کێ هاگاذاری کەرۆ چانیشا کە نزیکوو مەرذەینێ؟
پەی رۆ پێذا بیەیت چێش کەری؟
ئەگەر یۆ جە هاوڕێکات ئینا غەریبیەنە و دوورا جە کەسووکاریش چ راڤێژێڤش پنە مذەی و چ تەگبیرێڤش پەی کەری؟
دەڤاو دەرمانی سونەتی یا نەریتی لا شمە کامێنێ؟
There follows a randomly selected sample of transcribed responses from the corpus.
وەڵا لا ئێمەڤە فرەتەر چێڤێڤ کە ڤیرما جە قەذیمەنە وەختێڤ یۆ فرە مانیۆ کاری فرەش کەرذەن، وەختێڤ ملیەڤە یانە کرژ ماچی دابا چەیێڤ بوەروو. ئاچەیە فرە فرە ئێێ یۆی مسنۆڤە. ئێتر یۆ فرە مسیۆڤە پا چەیە. وەلی ئێسە ئاننە نەمەنەن ئیسە ماچا دابا قاڤێڤ بوەروو ئەڵبەتە قاڤە ساعبانێن کە ویشەڤە ئەجۆ ئێێ هینێذ کەرۆڤە ئێتر وەرمذ مەمەنۆ. تا ئێێێ هینێڤی بەشاشە بیەڤە. من وێم دلی ئاڤێنە وەختێڤ ملوو مەلێ کەروو دلێ ئاڤێنە فرە وەششا فرە حەزش پنە کەروو. یا رانە لڤەی، رانە لڤەیچ وەشا. یا مەسەڵەن تەماشدەو فیلمی کەروو. فیلێڤی وەش، ئانە فرە وەشا. ئێ ءیت عەر ئانە ئێ چەیێڤ بوەری قاڤێڤ بوەری ئەڵبەتە قاڤە چنی تیکێڤ کەیکێ بوەری فرە وەشا واقێحەن مسیێڤە. بلیەنە دلێ مڵکێڤی یا ئەگەر، یا بلی بلی یاگێڤە کە گوڵوو گوڵزار بۆ. ئانە فرە وەشا مسیەیەڤە واقێعەن
*واقێعەن موشکێڵەڤە جیدیەنە، چی دەورانوو ئێمەنە، فرە موشکێڵەڤە جیدیەنە ئەڵبەتە ئەرێ بنیاذم حەر تەسیرش هەنە سەروو تەبیعەتیەڤە بەڵام ئەگەر دەخالەتێ نەکەرا ویش ویش دەرمان کەرۆ. پسە زەلزەلەو تووفانی ئینێ هەر بیێنێ. ئتر پینیشا ویش دەرمانش کەرذەن زەمین هۆر وەلی چی ئەپێسە ماشینا کارخانێ من ئەجۆم ماشینێچ فرە فرە خرابتەرێنێ جە کارخانا ماشینەکێ فرە فرێنێ وە واقێعەن فرە فرە تەئسیرش هەن. *
عەز کەروو گرذێما وێموو زاڤڵێم گلێرێ بیمێڤێ کورذەسان وەشما بۆ، وەشما بۆ بیمێ پیشە ڤیس ساڵا چیەوڤەڵتەر گرذێما وەشما بۆ
*من نەلینا مەدرەسەو مامۆسام نەبیەنوو بەڵام لڤەینمێ نما وەنتەی لا مەڵەی. نماش فێرە کەرذێنمێ، نماش فێرێ کەرەذێنمێ و بەحزێڤ چێڤیش فێرێ کەرذێنمێ ئێتر خرابێ نەبیێنێ وەڵا. مامۆسیڤی فرە فەقیروو عال بێ. *
ئێ هیچ مامۆسەیڤم نەتاڤانش ژیانم ڤارۆ. ئێێ من ئەجۆم بەحزێڤ جارێ بە نێگاتیڤ مامۆساکێما مامۆاکێ کوردەسانی نێگاتیڤ ژیانوو منشا ڤاران مەسخەرەشاکەرذەنا چوون تاتەم موحەلم بیەن و ئێنتێزارشا بیەن ن زرێنگ بوو و زرێنگ نەبیەنا فرە مەسخەرەم دیەن جە مەدرەسەنە فرە ڤاتەنشا تۆ اتتەذ پێسنەنە مشۆ پاسنە بی تۆ تاتەذ چێمنەنە مشۆ پیسنە بی ئێتر ئێێێێ تەئسیرێڤی فرە مەنفیش سەروو زندگی منەڤە بیەن بەنەزەرۆ من مامۆسای موعەلێمێ وە ئەسڵەن هیچ موعەلێمێڤم وەشم نەسیان جە کورذەسانەنە بەڵام جە سوێدەنە یەکدانە موحەلم بێ من وەختێڤ ئامانێ سوێد هیچ بە حەیاتم کامیوترم نەذیەبێ ئێ بەڵام کە ئامانی سوید یەکەم جار لڤانێ کامپیوتێر کامپیوتێر چێڤێڤ بێ کە من حەز کەرێنێ بزانووش دماو ئانەیە حەز کەرێنیچ پرۆگرام بنڤیسوو پرۆگرامینگگ ئێ پرۆگرام بەرنامە نڤیسی ئێ عەز کەرێنێ بەرنامە نوش نوش نڤیسەی ئێت ئێێێێ مامۆسێڤی تەقریبەن بە تەمەن بێ جە دەبیرستانی بۆزۆرگساڵەنە کە من گەرەکم بێ دیپلۆم گێروو بلوو بە دیپلۆموو سوێدی بلوو دانشگا مامۆساوو ئێ ئێ هینی بەرنامە نڤیسەی بێ جە ئێبتێدایی تەرین بەرنامە نڤیسەی بەرنامەو سی سی پڵاس مامۆسێڤی بێ ئەنداز حال بێ مۆتەڤەجێ بێ کە من ئا مامۆسا مۆتەڤەجێ بێ کە من هیچ مەزانوو بەڵام خواستێڤی بێ ئەنداز گەورەم هەم گەرەکما بیاڤوو پنە ئێتر ئانە فرە ئێ یاڤەریش دانێ کە من حەتا چێڤانێ ئێبتێدایی تەرین چێڤانێ نەزانێنێ ئاذ یاڤەریش دانێ کە من ئا ئێبتێدایی تەرین چێڤا بزانوو وە فرە بە وەشروانە چنی من بەرخورذش کەرذ کە بی باعێسوو ئانەیە کە من بە ڕاسی حەز کەروو بەرانمە نڤیسەی
Each row of a tsv file represents a single audio clip, and contains the following information:
client_id - hashed UUID of a given user
audio_id - numeric id for audio file
audio_file - audio file name
duration_ms - duration of audio in milliseconds
prompt_id - numeric id for prompt
prompt - question for user
transcription - transcription of the audio response
votes - number of people that who approved a given transcript
age - age of the speaker1
gender - gender of the speaker1
language - language name
split - for data modelling, which subset of the data does this clip pertain to
char_per_sec - how many characters of transcription per second of audio
quality_tags - some automated assessment of the transcription--audio pair, separated by |
transcription-length - character per second under 3 characters per second
speech-rate - characters per second over 30 characters per second
short-audio - audio length under 2 seconds
long-audio - audio length over 5 minutes
non-allowed-script - transcription contains characters from a writing system not associated with the language
mixed-script-words - a single word contains characters from multiple writing systems
mixed-script-transcription - transcription spans multiple writing systems, but each word consistently uses only one
This dataset was partially funded by the Open Multilingual Speech Fund managed by Mozilla Common Voice.
This dataset is released under the Creative Commons Zero (CC-0) licence. By downloading this data you agree to not determine the identity of speakers in the dataset.
For a full list of age, gender, and accent options, see the demographics spec. These will only be reported if the speaker opted in to provide that information. ↩ ↩2