License:
SURZHYK-TERMS-0.1.0
Steward:
CommunityDataset ID:
cmuv814nm0040ns08kyzpnry4
Release Date: 10/5/2026
Format: WAV, TSV, MD
Size: 240.92 MB
This dataset contains 3,117 prompted voice recordings, totaling approximately 3.09 hours, from 16 pseudonymous contributor accounts. The speech includes Ukrainian, Russian, and mixed forms associated with Surzhyk. Audio is provided as mono 16 kHz, 16-bit PCM WAV files, paired with transcripts, stable clip andsplits are included in this preliminary release; future versions may add reviewed transcripts and annotations.
Licensing
Surzhyk Read Speech Dataset Use Terms 0.1.0
https://github.com/ilyash819/surzhyk-corpus/blob/07cc335255d3cb275c526bf74e561b2bdd366ee5/TERMS.mdRestrictions/Special Constraints
The recordings and original prompts may be used for speech recognition, machine learning research and development, and linguistic research, subject to TERMS.md. Syrzyk source text remains separately licensed under CC BY 4.0.
Forbidden Usage
You must not use the recordings to identify or attempt to identify speakers, link them to real-world identities, or enroll them in biometric identification systems. Voice cloning and training or using models to generate speech that imitates contributors are prohibited.
Ethical Review
Participants were shown a consent notice explaining the dataset’s purpose, potential publication of recordings and transcripts, and their use for speech recognition, machine learning, and linguistic research. They were asked to confirm that they were at least 18 years old and participating voluntarily. The notice explained that complete anonymity cannot be guaranteed because voices may be recognizable. Public files use pseudonymous speaker IDs rather than account identifiers. Dataset terms prohibit speaker identification, biometric enrollment, and voice cloning.
Intended Use
This dataset is intended for training and evaluating automatic speech recognition systems for Ukrainian, Russian, and mixed speech associated with Surzhyk, and for linguistic research on language mixing. Transcripts should be checked against the audio before use as evaluation references.