Release Date: 9/7/2026
Format: WAV, CSV
Size: 17.46 GB
A Polish-language dataset for the detection of synthetic speech. Most recordings belong to a pair: a synthetic recording and a human recording of the same text, so that detection models are trained and evaluated on differences in signal rather than in content. About a quarter of the recordings have no counterpart. Three generation methods are covered: text-to-speech with catalogue voices, voice cloning from reference recordings, and retrieval-based voice conversion. Its human recordings come from six publicly available Polish speech corpora. Most of the synthetic ones were generated for this dataset with commercial and open-source systems. The rest are taken from the MLAAD dataset. Every recording is labelled at file level as human or machine-generated, and the index records the system that produced it and the corpus it comes from.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
Use of this dataset as a whole is limited to non-commercial research. The subsets MAILABS, nEMO and mozilla carry the same limit on their own. It does not apply to fleurs (CC BY-SA 4.0) or generated (CC BY 4.0), whose licenses permit commercial use. It also does not apply to material dedicated to the public domain under CC0 1.0, which remains in the public domain wherever it appears. In the mozilla subset that is the Common Voice, Darkman 1.0 and Gosia 1.0 recordings and their transcripts, which are the larger part of it. What stays non-commercial there is the compilation and the synthetic recordings created for this dataset. The Common Voice audio of the mozilla subset — the Common Voice recordings and every synthetic recording made from them — carries two further conditions, listed under Forbidden Usage. It may not be used to train speech-synthesis or voice-conversion models, and it may not be re-hosted outside this platform. Both were offered to Mozilla in the request for permission to publish the Common Voice portion, and both are accepted by every person who downloads this dataset from this platform. Neither extends to the Darkman 1.0 and Gosia 1.0 recordings in the same subset. Both extend to every synthetic recording made from a Common Voice recording, whatever voice it carries.
Share
Forbidden Usage
It is forbidden to attempt to determine the identity of any speaker in this dataset. It is forbidden to use this dataset in any way likely to harm the individuals whose voices it contains, and that includes any such use of the synthetic recordings that imitate a real speaker's voice. This covers impersonation, fraud, and any use of a cloned voice to represent a person as having said something they did not say. It is forbidden to use the Common Voice audio of the mozilla subset — the Common Voice recordings and every synthetic recording made from them — to train, fine-tune or improve any speech-synthesis or voice-conversion model. It is forbidden to re-host or re-share that same audio outside Mozilla Data Collective, which is the single point of distribution for it, unless the dataset page states that distribution has moved elsewhere. These last two conditions do not extend to the Darkman 1.0 and Gosia 1.0 recordings in the same subset, which their publisher released under CC0 1.0 without them. They do extend to every synthetic recording made from a Common Voice recording, including one converted into a Darkman 1.0 or Gosia 1.0 voice.
Ethical Review
The dataset was assembled for scientific research, with the safeguards that Article 89 GDPR, read with Recital 159, requires of research processing. The platform does not require a separate data-protection review, and the safeguards below were designed and applied for this dataset. Speakers appear only under pseudonymous identifiers, and no attempt to determine their identity is permitted under the terms of access. Access is granted on request, so that a withdrawal of consent can be passed on to everyone holding the data. Recordings withdrawn by their contributors in any Common Voice release are removed together with every synthetic recording derived from them. The external_ref and group_id columns make this traceable to the individual file. The age field of Common Voice is not published. Voice cloning was applied to 72 speakers: 59 from Common Voice, 9 nEMO actors, 2 M-AILABS readers, and the 2 donors of the Darkman 1.0 and Gosia 1.0 corpora, whose recordings were published for exactly this kind of use. No voice was cloned from a person known to the publisher to be a public figure. Of the 59 Common Voice speakers, 31 declared an age, and every one of those declarations falls in an adult band. The remaining 28 declared none, and an indirect safeguard covers them: voices whose median fundamental frequency fell in the range typical of children were kept out of the cloning pool. That is a screen on pitch, not a verification of age. No cloning or voice-conversion model is distributed with this dataset. The conversion models were fine-tuned on the M-AILABS, Darkman 1.0 and Gosia 1.0 corpora alone, whose licenses permit that use, starting from publicly released base weights whose authors do not document their training data. Every synthetic recording is marked as machine-generated at file level, with its generating system recorded alongside.
Intended Use
Research on the detection of synthetic speech in Polish: training and evaluating detection models, benchmarking across generation methods, and studying generalisation to unseen synthesis systems. The paired structure also supports work on the acoustic differences between human and synthetic speech. Use is limited to non-commercial research.
101 175 recordings, 142.3 hours — 71.7 hours of human speech and 70.6 hours of synthetic speech, from 1 279 human speaker identifiers and 18 synthesis and voice-conversion systems, ten of them commercial services.
| Subset | Source corpus | Human recordings | Synthetic recordings | Subset license |
|---|---|---|---|---|
MAILABS | M-AILABS (Polish) + MLAAD | 6 239 | 10 519 | CC BY-NC 4.0 |
nEMO | nEMO (emotional speech) | 4 481 | 8 159 | CC BY-NC-SA 4.0 |
fleurs | FLEURS (Polish) | 3 937 | 3 576 | CC BY-SA 4.0 |
mozilla | Common Voice, Darkman 1.0, Gosia 1.0 | 34 203 | 23 811 | CC BY-NC-SA 4.0 |
generated | Sentences generated with a language model | — | 6 250 | CC BY 4.0 |
| Total | 48 860 | 52 315 |
FLEURS publishes no speaker identifiers, so all its human recordings are attributed to two placeholders, fle-unknown_male and fle-unknown_female. The speaker count above is therefore a lower bound.
The commercial systems are Google Chirp 3 HD, ElevenLabs in three models (Multilingual v2, Eleven v3 and Flash v2.5), MiniMax Speech 2.6 HD, Speechify Simba Multilingual, Microsoft Azure AI Speech, OpenAI GPT-4o Mini TTS, and Amazon Polly in its Standard and Generative engines. The rest are RVC voice conversion, run by the dataset authors, and seven systems whose recordings come from MLAAD: Coqui XTTS v1 and v2, Microsoft Edge TTS, Llasa 1B Multilingual, OuteTTS, VITS MAI PL and WhisperSpeech.
The archive extracts to a single directory holding the LICENSE and NOTICE files and one directory per evaluation split, each with an audio/ directory of recordings and a metadata.csv index. The index is pipe-delimited, UTF-8 encoded and written with Unix line endings, and audio_path is relative to the split directory. Every row carries:
| Column | Meaning |
|---|---|
file_id, audio_path | identifier of the recording, and its location within the split directory |
split | the split the row belongs to, so the eight indexes can be concatenated |
dataset | source subset: MAILABS, nEMO, fleurs, mozilla, generated |
type, label | human / tts, and 0 for a human recording, 1 for a synthetic one |
speaker_id | the speaker of a human recording, or the voice of a synthetic one |
gender | male, female or unknown. See the note on gender labels below |
normalized_text | the spoken text, lower-cased and with punctuation removed. Normally identical within a pair. See the note on transcripts below |
model_id, provider | the system that produced the recording, and the provider whose terms apply to its output. Both are empty for human recordings. In provider, RVC marks the voice conversion run by the dataset authors and MLAAD marks recordings taken from that dataset. See the note on model identifiers below |
system_type | human, commercial or open-source: whether the recording is a human one, or came from a commercial service or an open-source system. The terms of that system govern what may be done with the audio |
external_ref | for the Mozilla Common Voice material, the name of the source clip in that corpus. Carried both by the human clip and by every synthetic recording in its group |
group_id | ties a human recording to every synthetic recording made from it. Where a group holds no human recording, is_paired is False throughout it |
is_cloned | True where the synthetic voice was cloned from reference recordings of a specific speaker. Voice-conversion recordings carry False — their target voice is named by speaker_id |
is_paired | whether the group holds both a human and a synthetic recording |
Gender labels. For the corpora that publish a gender field, the value is theirs. For Common Voice it is mixed: most speakers declared a gender, and for the rest it was inferred from the fundamental frequency of their clips, at 97.3 % agreement with the declared value. The column does not distinguish the two, so treat gender in the Common Voice material as an attribute of the recording rather than a statement about the person.
Model identifiers. model_id takes 90 values across the 18 systems, because most of them name the voice and the generation date alongside the system. Group on the leading part of the value when counting systems.
Transcripts. Normalisation removes sentence punctuation but is not exhaustive, and a small number of rows keep a quotation mark, apostrophe, slash or per-cent sign. Every row carries the text that was actually read or synthesised, which is why in a few groups the two sides of a pair differ in how a numeral or an apostrophe was written. Recover pairs through group_id rather than through text equality.
All recordings are mono WAV, 16-bit PCM. Sample rates are not uniform. Each recording keeps the rate of the corpus it comes from or of the system that produced it, and the files are published exactly as they were returned, with no resampling.
Resample every recording to a common rate before training or evaluating.
The splits are not an arbitrary partition. Two are for fitting a model and six are test sets, each holding one experimental condition so that results can be read per set without filtering:
| Split | Recordings | Condition |
|---|---|---|
train | 64 075 | training material |
val | 7 228 | validation, for threshold selection |
test_a_standard | 6 704 | catalogue text-to-speech, speakers seen in training |
test_b_voice_cloning | 4 654 | voice cloning, speakers seen in training |
test_c_rvc | 8 000 | voice conversion, speakers seen in training |
test_d_voice_cloning_unseen | 2 514 | voice cloning, speakers absent from training |
test_e_rvc_seen | 4 000 | voice conversion, seen speakers, previously unused clips |
test_f_rvc_unseen | 4 000 | voice conversion, speakers absent from training |
Test-E and Test-F were drawn in a single pass, with the same procedure and the same target voices, so that speaker novelty is the manipulated variable. They are the matched pair for measuring its cost. Test-B and Test-D pose the same question for voice cloning, but are matched less tightly, coming from different Common Voice releases. The seen and unseen contrast describes each split as a whole rather than every row in it.
No voice-conversion material is present in train or val, so Test-C, Test-E and Test-F also measure generalisation to an unseen conversion system. Pairing is complete in Test-B through Test-F. Test-A is the only test set with unpaired material, human and synthetic alike. The pairing is carried by group_id, and is_paired states it per row.
The dataset is a collection of separately licensed subsets, released as a whole under CC BY-NC-SA 4.0, which is the value in the platform's license field. The non-commercial condition comes from the nEMO corpus and the MLAAD dataset, the share-alike condition from nEMO alone. The fleurs subset is released under CC BY-SA 4.0 and generated under CC BY 4.0. Material sourced from Mozilla Common Voice, Darkman 1.0 and Gosia 1.0 is dedicated to the public domain under CC0 1.0 and remains there — the non-commercial condition is not extended to it.
The full structure, the per-subset breakdown and the attributions required by the source corpora are in the LICENSE and NOTICE files inside the archive.
The Common Voice portion consists of clips selected from the validated set of the Polish corpus, releases 24.0 and 26.0, restructured for a detection task and published together with synthetic counterparts generated from the same transcripts. Some of those counterparts are voice clones, made from the speakers' own validated clips, because detection of cloned speech cannot be measured without examples of it.
This is not an alternative distribution channel for Common Voice. The canonical source of the corpus remains Mozilla Data Collective, and this dataset is published with Mozilla's permission, granted on the condition that the Common Voice audio is distributed only from this platform. Recordings whose contributors withdraw consent are removed on each Common Voice release, together with every synthetic recording derived from them. The safeguards behind that commitment are set out in the ethical review section of this datasheet.