License:
CC-BY-4.0
Steward:
Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOEDataset ID:
cmuzsrkif007m07nxz97doh8l
Release Date: 10/8/2026
Format: WAV
Size: 6.94 MB
UyghurDF is a Uyghur speech deepfake dataset and benchmark designed for evaluating speech deepfake detection and cross-generator generalization. The complete dataset contains 315,629 utterances (577.56 hours) from 1,874 speakers and covers 18 generation systems spanning text-to-speech, voice conversion, neural vocoder reconstruction, and neural codec reconstruction. This release is UyghurDF Preview v0.1, a small representative subset containing 47 WAV files, including 3 bona fide and 44 spoof samples. It is provided for inspecting the data format, metadata organization, and representative audio characteristics. This preview is not the complete benchmark release and should not be used to reproduce the full benchmark results. The full dataset is planned for release after acceptance of the associated paper.
Licensing
Creative Commons Attribution 4.0 International (CC-BY-4.0)
https://spdx.org/licenses/CC-BY-4.0.htmlRestrictions/Special Constraints
UyghurDF is released under CC BY 4.0. Materials originating from THUYG-20 remain subject to the applicable requirements of the Apache License 2.0, including preservation of the license and relevant copyright, attribution, and NOTICE information where applicable. Materials originating from Mozilla Common Voice remain subject to the applicable Common Voice and Mozilla Data Collective terms, including applicable access and redistribution requirements. Users must not attempt to determine the identity of Common Voice speakers. Please refer to the accompanying LICENSE, THIRD_PARTY_NOTICES.md, and LICENSES/ for details.
Forbidden Usage
Users must not attempt to identify, re-identify, or infer the identity of speakers represented in the dataset. Materials originating from Mozilla Common Voice must not be used in violation of the applicable Common Voice or Mozilla Data Collective terms. Users must comply with all applicable licenses, source-specific terms, and legal requirements. The dataset must not be used for discriminatory purposes, including ethnic profiling, targeted surveillance, persecution, or other activities that violate the rights, privacy, or safety of individuals or communities.
Ethical Review
UyghurDF was constructed using existing speech corpora and generated speech, and no new human-subject speech data were collected specifically for this dataset. Source materials are reused in accordance with their applicable licenses, permissions, and platform terms. The dataset is intended for speech deepfake detection and defensive research, and users are required to respect applicable privacy, licensing, and source-specific restrictions.
Intended Use
UyghurDF is intended for research and benchmarking on Uyghur speech deepfake detection, speech anti-spoofing, cross-generator generalization, low-resource speech forensics, and generator attribution. It can be used to develop and evaluate models for distinguishing bona fide from spoofed speech, identifying the generation system responsible for spoofed speech, and studying robustness to unseen speech generation systems.
UyghurDF is a Uyghur speech deepfake dataset designed for research on speech deepfake detection, speech anti-spoofing, generator attribution, and cross-generator generalization.
The dataset contains bona fide Uyghur speech and spoofed speech produced using 18 speech generation or reconstruction systems spanning text-to-speech (TTS), voice conversion (VC), neural vocoder resynthesis, and neural codec reconstruction.
The full UyghurDF dataset contains:
All audio in the final dataset is standardized to:
The current release on Mozilla Data Collective is a preview subset of UyghurDF rather than the complete dataset.
This preview is provided to demonstrate the dataset structure, metadata, audio format, and the range of spoof generation systems included in UyghurDF.
The current preview contains 47 WAV files:
The bona fide samples included in the current preview originate from THUYG-20.
The construction procedure and statistics described below refer to the complete UyghurDF dataset represented by this preview.
The complete UyghurDF dataset will be uploaded and made publicly available after the associated paper is formally accepted.
The primary language of the dataset is Uyghur.
Associated transcripts are primarily written using the Uyghur Arabic script.
Bona fide speech in the full UyghurDF dataset is collected from two existing Uyghur speech resources:
For Common Voice, client_id is used as the speaker identifier.
For THUYG-20, the original speaker identifiers provided by the corpus are used.
For Common Voice samples with missing gender information, gender metadata was completed using available metadata where possible or through manual annotation.
Audio collected from Common Voice Uyghur and THUYG-20 was first converted to a common technical format:
Speaker identity, gender information, and transcript information were then organized for subsequent dataset construction.
The bona fide data was partitioned into Train, Development, and Test sets at the speaker level.
All utterances from the same natural speaker are assigned to only one partition. Therefore, natural speakers do not overlap between the Train, Development, and Test sets.
The final bona fide data contains 157,695 utterances:
Spoof generation was performed only after the speaker-level partitioning was completed.
Generated samples inherit the partition associated with their corresponding source speech so that source-speaker information does not cross dataset partitions.
Spoof speech was constructed using three main generation procedures:
For analysis and benchmarking, speech reconstruction is further divided into:
Uyghur text was collected from transcripts associated with Common Voice Uyghur and THUYG-20.
Before synthesis, the text underwent basic cleaning, including removal of empty entries and duplicate text.
The resulting Uyghur text was then provided to different TTS systems to generate spoofed speech.
Some TTS systems use a fixed-speaker or fixed-voice configuration. Consequently, these systems contain fewer generated samples than many voice conversion and reconstruction systems.
Voice conversion uses speech from Common Voice Uyghur and THUYG-20 to provide linguistic content and target-speaker characteristics.
For each VC pair:
This prevents natural-speaker information from crossing partition boundaries during voice conversion.
VC pairing was designed to balance four source-to-target gender directions:
The four directions were sampled in approximately equal proportions during data construction. This provides balanced coverage of both same-gender and cross-gender conversion while also maintaining approximately balanced male and female target voices.
Speech reconstruction takes bona fide speech as input and reconstructs it using either a neural vocoder or a neural audio codec.
Source utterances were sampled with consideration of both source corpus and speaker gender in order to maintain reasonable coverage of:
For benchmarking purposes, reconstruction systems are divided into:
UyghurDF contains spoofed speech produced by 18 generation or reconstruction systems across four categories.
To evaluate generalization to previously unseen speech generation methods, the 18 generation systems in UyghurDF are divided into Seen and Unseen generators.
Seen generators are represented in the Train, Development, and Test partitions. These systems can therefore be observed during detector training and model selection.
Unseen generators are completely excluded from the Train and Development partitions and appear only in the Test partition. They are used to evaluate whether a detector can generalize to generation methods that were not observed during training.
| Category | Generation System | Protocol | Available Splits |
|---|---|---|---|
| TTS | MMS-TTS-Uyghur | Seen | Train / Dev / Test |
| TTS | MMS-TTS-Uyghur-UQSpeech | Seen | Train / Dev / Test |
| TTS | Coqui-VITS-Uyghur | Seen | Train / Dev / Test |
| TTS | OmniVoice | Seen | Train / Dev / Test |
| TTS | XFYun | Unseen | Test only |
| VC | Seed-VC | Seen | Train / Dev / Test |
| VC | FreeVC | Seen | Train / Dev / Test |
| VC | OpenVoice | Seen | Train / Dev / Test |
| VC | FACodec | Seen | Train / Dev / Test |
| VC | MKL-VC | Seen | Train / Dev / Test |
| VC | X-VC | Seen | Train / Dev / Test |
| VC | GenVC | Unseen | Test only |
| VC | EZ-VC | Unseen | Test only |
| VC | Vevo2 | Unseen | Test only |
| Neural Vocoder Resynthesis | BigVGAN | Seen | Train / Dev / Test |
| Neural Vocoder Resynthesis | HiFi-GAN | Seen | Train / Dev / Test |
| Neural Codec Reconstruction | EnCodec | Seen | Train / Dev / Test |
| Neural Codec Reconstruction | DAC | Unseen | Test only |
In total, UyghurDF contains:
The spoof portion of the Test set contains:
The five Unseen generators cover multiple generation paradigms:
This design allows UyghurDF to support two complementary evaluation scenarios:
The second setting is intended to measure cross-generator generalization rather than only performance on known spoofing methods.
Generated and processed audio was screened for abnormal or near-silent samples.
Each utterance was divided into non-overlapping 20 ms frames.
An utterance was flagged by the silence-screening procedure when either of the following conditions was satisfied:
This procedure was used to identify failed or near-silent generation outputs.
In addition to automatic screening, manual quality inspection was conducted for the generated speech.
For each generation system, five samples were randomly selected for manual inspection, with both male and female speech included where applicable.
The inspection was conducted by native Uyghur speakers.
Depending on the generation method, reviewers were provided with relevant information such as:
The inspection focused primarily on whether:
The inspection showed that generation quality varies across systems.
Some generated utterances contain pronunciation or linguistic-content differences relative to the corresponding source or expected speech.
These differences originate from the behavior and limitations of the generation systems themselves and are retained as part of the variation among spoofing methods represented in the dataset.
Because different speech generation systems produce audio using different sampling rates, channel configurations, and encoding formats, a final global standardization step was applied after bona fide and spoof speech were combined.
All final audio files were standardized to:
During this final standardization stage, 101,484 audio files required format conversion.
All 101,484 files were successfully converted, with no conversion failures.
The full UyghurDF dataset contains 315,629 audio files.
| Split | Bona fide | Spoof | Total | Duration |
|---|---|---|---|---|
| Train | 114,868 | 114,940 | 229,808 | 412.75 h |
| Development | 11,346 | 11,494 | 22,840 | 43.54 h |
| Test | 31,481 | 31,500 | 62,981 | 121.27 h |
| Total | 157,695 | 157,934 | 315,629 | 577.56 h |
The complete dataset contains:
The number of bona fide and spoof utterances is kept close to a 1:1 ratio.
The Train, Development, and Test partitions remain strictly disjoint with respect to natural speakers.
Metadata is provided to describe the audio samples and their roles in the benchmark.
Depending on the sample type, metadata may include information such as:
For generated speech, additional information related to the generation procedure may also be recorded where applicable.
UyghurDF is intended for research and benchmarking in areas including:
The dataset is primarily intended for defensive and academic research.
UyghurDF currently focuses on fully spoofed utterances.
Partial speech manipulation, local editing, and splicing attacks are not included in the current version.
Although the dataset includes 18 generation systems, these systems cannot represent all existing or future speech synthesis, voice conversion, neural vocoder, or neural codec technologies.
Generation quality also varies across systems. Some systems may introduce pronunciation errors, linguistic-content differences, or synthesis artifacts.
Performance measured on UyghurDF should therefore not be interpreted as guaranteed performance against all possible speech deepfake generation methods, languages, recording environments, or deployment conditions.
UyghurDF was constructed using existing speech corpora together with generated and reconstructed speech.
No new human-subject speech recordings were collected specifically for the construction of UyghurDF.
Source materials remain subject to their respective licenses, permissions, and platform terms.
Materials originating from Mozilla Common Voice and THUYG-20 should be used in accordance with the applicable source-specific terms and licenses.
Users of UyghurDF are responsible for complying with all applicable source-specific licensing, redistribution, privacy, and usage requirements.