Release Date: 9/17/2026
Format: MP3
Size: 228.47 MB
A collection of spontaneous responses to questions in Scots (Scots).
Restrictions/Special Constraints
None provided.
Forbidden Usage
It is forbidden to attempt to determine the identity of speakers in the Common Voice datasets. It is forbidden to re-host or re-share this dataset.
Intended Use
This dataset is intended to be used for training and evaluating automatic speech recognition (ASR) models. It may also be used for applications relating to computer-aided language learning (CALL) and language or heritage revitalisation.
sco)This datasheet is for sps-corpus-5.0-2026-09-11 of the Mozilla Common Voice Spontaneous Speech dataset for Scots [Scots - sco]. The dataset contains 715 clips representing 11.17 hours of recorded speech (10.68 hours validated) from 21 speakers.
Scots, a sister language of English spoken throughout Scotland, has a long history. It arose from northern English dialects around the 14th century and spread east and northwards, supplanting the indigenous Gaelic language, and developing into a socially and politically high status language with spoken and written norms distinct from those in England. From the 17th century onwards, political, religious and social events led to a loss of status, and thus a shrinking of the domains in which Scots was used, but while English norms replaced Scots in writing, spoken Scots continued to be used. Present day Scots is characterised by the Scots Linguistic Continuum with Standard Scottish English – generally described as being close to Standard English but with an overlay of distinctly Scottish sounds – at one end and Broad Scots – much further from Standard English with its own words, sounds and sentence structures – at the other. In terms of the social profiles, Scottish Standard English is spoken by middle class speakers and in more formal situations such as in schools, while Broad Scots is spoken by working class speakers and in informal situations such as with family and friends. Speakers may styleshift up and down the continuum according to, amongst others, interlocutor and context. At the Broad Scots end of the continuum, there is significant geographic diversity, where, for example, speakers in Glasgow sound very different to speakers in Aberdeen. The speakers in these recordings are from a range of geographic locations, and align more with the Broad Scots end of the continuum.
The dataset clips are categorised by transcription status and training-set assignment. The following tables summarise the distribution.
| Bucket | Clips | % |
|---|---|---|
| Transcribed & Validated | 680 | 95.1% |
| Transcribed & Pending | 0 | 0.0% |
| Not transcribed | 35 | 4.9% |
| Bucket | Clips | % |
|---|---|---|
| Train | 464 | 64.9% |
| Dev | 83 | 11.6% |
| Test | 133 | 18.6% |
| Unassigned | 35 | 4.9% |
Training split coverage: 680 of 680 transcribed & validated clips (100.0%)
The transcription system uses general Latin script.
Prompts: 47
Duration: 40234608[ms]
Avg. Transcription Len: 725
Avg. Duration: 56.27[s]
Valid Duration: 38478.56[s]
Total hours: 11.18[h]
Valid hours: 10.69[h]
| Bucket | Clips | % |
|---|---|---|
| Validated | 680 | 100.0% |
| Pending | 0 | 0.0% |
| Edited | 190 | 27.9% |
Present day Scots has no written standard and orthographic conventions vary both within and between the different dialects being represented. For example, can’t may be cannae or canny in Edinburgh, but canna in Aberdeen. For these transcriptions, we have followed protocols documented in previous research e.g. https://scotssyntaxatlas.ac.uk
There follows a randomly selected sample of questions used in the corpus.
What’s the most expensive thing you’ve ever bought?
What’s the most stressful part of your work?
What’s the worst technological advancement in recent years?
What was your favourite subject in school and why?
Who won in a fight with your siblings?
There follows a randomly selected sample of transcribed responses from the corpus.
What’s the best night out you’ve ever had? Och, [inc]. I don’t know, there’s been loads of them. [laugh] When we were- when I was younger, I used to go to a nightclub called Jackie O’s and Bently’s they were in Kirkcaldy, and that was the disco, but it was great fun, and I used to enjoy going on a Wednesday night, which was older music, I liked older music, and then, on a Friday night, or a Saturday night, and people used to come fae all round, so yeah, I think, I think that’s probably- well I had lots of good nights out, but I think they were the best nights that we’ve ever had, or I’ve ever had.
*Favourite book. Well, when I was younger I always used to love Oor Wullie. I'd get it annually every Christmas, eh, in the eighties- in the seventies and eighties. Eh, used to love it. Eh, it was a’ aboot, know, a wee guy. Used to go aboot with his pals Boab and Wee Eck. And he would g- get into trouble jumping over dykes and a’ that and then running after the big- the big sergeant- would be chasing him [inc]. Just some of the things were amazing. Different stories. Wee captures on every wee page and it was- know. And it led to kinda The Broons and a’ that and it was really good. Eh, it was a’ aboot like Dundee. I think it was maistly based in Dundee kinda patter. Oh and, eh, things like that and then I used to love the Dandy which was just something [inc], know the- the big- big, eh, heavy man. Desperate Dan. Used to love a’ that. It was great. Desperate Dan and- and then the- the Beano too with Gnasher. And eh, the guy- what’s his name. The- wee- Dennis the Menace. And he was with his black and white stripy jumpers. Brilliant. And we had them when we were wee and a’ that too, you know? Felt like wee comic guys [laugh]. *
*The skill I would maist like to learn would be, just to be a mechanic, that’s what I would like to d- to ken athing there is to do with cars, athing there is to do with motors ‘cause I seem to aie just be spending millions of money, millions and millions of money on my motor, and eh, there’s aie stuff going wrong with it, and then you’ve got to book it tak it to the garage, and then they charge you through the nose for a’ the work that needs done to it. And I just think, if I could do this myself, I’d save millions of time and millions of money, so that’s what I would like to- that’s the skill I would like is to- to ken athing that there is to ken aboot motors. Basically, be a mechanic. *
*One- one of the things that eh is most stressful about my job is that we are dealing with, eh, clients that are, eh have got a lot of substance misuse issues or alcohol issues, mental health issues, and a lot of them can be challenging in the way that they speak to you. They can shout, they can swear, eh, and eh sometimes it’s really really difficult to try and- I don’t know quite- get through to them or listen to them. Sometimes people deliberately shout and respond deliberately in that manner because they don’t want you to challenge eh, maybe, you know, or try and find them out because you’re maybe asking a reason why they haven’t attended or eh and some people I feel in our work have eh learnt that by shouting the loudest they’ll probably get, eh, more, they’ll probably get what they want, because they’re making a scene, eh but I think that’s what’s maist stressful. If someone’s- if you’re dealing with someone and you’re trying to speak to them and reason with them, and sometimes you can’t reason with them, and that’s no a reflection on me or yourself and I think that’s what you’ve got to eh take away eh from it, but sometimes it’s- it can really just upset your day sometimes, and, but I think as well in saying that you know th- we know the type of people that we’re dealing with and we know that that can change in an instant. Their- their- their hum-their demeanour, their language can change in a minute, eh and their tempers can go fae zero to a hundred eh in a second. So, eh, it’s- it’s really frustrating when you feel like they’re not listening, eh, you cannae help them, eh and it’s because really they dinna want to help themselves. *
*Eh the most expensive thing that I’ve ever bought’s a holiday. Eh and eh, yeah [laugh] eh not- eh anything else really, I could say, but yeah, that’s the most expensive thing I’ve ever bought, eh, yeah. *
Each row of a tsv file represents a single audio clip, and contains the following information:
client_id - hashed UUID of a given user
audio_id - numeric id for audio file
audio_file - audio file name
duration_ms - duration of audio in milliseconds
prompt_id - numeric id for prompt
prompt - question for user
transcription - transcription of the audio response
votes - number of people that who approved a given transcript
age - age of the speaker1
gender - gender of the speaker1
language - language name
split - for data modelling, which subset of the data does this clip pertain to
char_per_sec - how many characters of transcription per second of audio
quality_tags - some automated assessment of the transcription--audio pair, separated by |
transcription-length - character per second under 3 characters per second
speech-rate - characters per second over 30 characters per second
short-audio - audio length under 2 seconds
long-audio - audio length over 5 minutes
non-allowed-script - transcription contains characters from a writing system not associated with the language
mixed-script-words - a single word contains characters from multiple writing systems
mixed-script-transcription - transcription spans multiple writing systems, but each word consistently uses only one
Jennifer Smith <[email protected]>
This dataset was partially funded by the Open Multilingual Speech Fund managed by Mozilla Common Voice.
This dataset is released under the Creative Commons Zero (CC-0) licence. By downloading this data you agree to not determine the identity of speakers in the dataset.
For a full list of age, gender, and accent options, see the demographics spec. These will only be reported if the speaker opted in to provide that information. ↩ ↩2