License:
GPL-3.0
Steward:
CommunityDataset ID:
cmuy24yjx00vo07mf3azmkkf2
Release Date: 10/7/2026
Format: WAV, JSONL
Size: 550.01 MB
2,233 synthetic single-speaker utterances (5.78 h) of informal Spanish–English code-switched conversation. Text was generated locally with an open LLM (gpt-oss-20b) and voiced with Qwen3-TTS, using voices designed from text descriptions only; no recording of a real person is included. The datasheet also describes the voice conversion that was applied to the copy used for training in the *Lost in Transcription — Spanish–English* competition (MDC/DrivenData, 2026); at MDC's request, that voice-converted audio is not distributed.
Licensing
GNU General Public License v3.0 or later (GPL-3.0)
https://spdx.org/licenses/GPL-3.0-or-later.htmlRestrictions/Special Constraints
No restrictions beyond GPL-3.0-or-later. Recipients must keep the licence and the NOTICE file. This release contains no audio from the Bangor Miami corpus. Audio produced with the voice-conversion recipe described in the datasheet is derived from that corpus and subject to its terms, including BangorTalk's condition that the material may not be used to train an AI unless all of that model's training data is made publicly available.
Forbidden Usage
None beyond the licence. Request (not a licence condition): please do not use this dataset, or the recipe in the datasheet, for speaker identification, voice cloning or impersonation of the Bangor Miami participants, or any other biometric purpose, and do not try to re-identify corpus participants.
Ethical Review
(1) Nature of the dataset: fully synthetic speech. No human recordings were made, all words are machine-generated, and every voice was designed from a text description with Qwen3-TTS VoiceDesign; no recording of a real person and no Bangor Miami audio is included. (2) Change made at MDC's request (review of 5 October 2026): an earlier file of this upload contained the same utterances after voice conversion (kNN-VC) toward 37 adult speakers of the Bangor Miami corpus. That version was used to train our competition models; it has been replaced by this version and is not distributed. The conversion is described in the datasheet ("Voice conversion in the competition training copy"). (3) Safeguards in the recipe: only adult target speakers (utterances converted toward the 2 child speakers that occurred were removed); target speakers are referenced only by opaque IDs (T01–T37) and the mapping to corpus speaker codes is not published; all transcripts are synthetic.
Intended Use
Training and evaluation of ASR for Spanish–English code-switched conversational speech; research on synthetic data for low-resource code-switching ASR.
Created by participant inyongkhafid for the Lost in Transcription — Spanish–English challenge (Mozilla Data Collective / DrivenData, 2026) and
published on MDC because the competition requires external training data to be published on MDC. License: GPL-3.0-or-later.
This version contains no voices of real people. All voices were designed from text descriptions with a TTS model. At MDC's request (review of 5 October 2026), the voice-converted copy that was used for training in the competition is not distributed. It is described below in "Voice conversion in the competition training copy (not distributed)".
DS-3 is a fully synthetic corpus of informal Spanish–English code-switched speech. 86 % of utterances contain both Spanish and English function words. It was built to reduce the mismatch between the Bangor Miami corpus (mostly English, little intra-utterance switching) and conversational voice-note style speech. A voice-converted copy of it (see below) was used, together with the Bangor Miami corpus, to fine-tune Whisper large-v3, Parakeet-TDT-0.6B-v3, Cohere Transcribe 03-2026 and Qwen3-Omni-30B-A3B in the Spanish–English submission that placed 1st on the competition's private leaderboard.
target_voice (used only by the voice-conversion recipe): 37 opaque IDs, one per adult Bangor Miami speaker (sastre 869 · herring 568 · zeledon 422 · maria 374 utterances by corpus family)| Step | What | Tool / model (licence) | Kept |
|---|---|---|---|
| 1. Text | 360 conversations × 12 turns were requested (informal es-en code-switching); 2,360 well-formed turns were kept. Each conversation has a short persona/topic plan; turn-level language mixing targets aggregate statistics of conversational code-switching (Spanish share 35–55 %, ≥ 70 % mixed turns, median 1–4 switch points, 15–40 words). Few-shot examples were written by hand. | openai/gpt-oss-20b (Apache-2.0), GGUF ggml-org/gpt-oss-20b-GGUF MXFP4 (sha256 27cd6c432c76…), served locally with llama.cpp (MIT) | 2,360 turns |
| 2. TTS | One clip per turn and speaker. For each persona, VoiceDesign creates a reference clip from a text description of the voice only (for example "A middle-aged bilingual Latino man from Miami, warm relaxed timbre, casual and fast conversational speech"); the Base model then clones that synthetic reference for every turn of the persona. No recording of a real person is used. Round-trip QC: an in-house ASR model transcribes each clip; clips with WER > 50 % are removed. | Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign + -Base (Apache-2.0) | 2,354 (QC corpus WER 5.35 %) |
| Selection | This release keeps the 2,233 utterances that make up the competition training copy (the others were removed by the QC of the voice conversion and by the child-voice filter, see below). Audio and text are exactly the step-2 output. | own code | 2,233 |
| Text statistics (step 1) and the timing targets of the acoustic-difficulty step of the recipe below were taken from aggregate, clip-level statistics of the competition development set (language-mix ratios; speech rate, voiced ratio, duration). No audio, transcripts or excerpts of competition data are contained in DS-3 (verified: 0 eight-word sequences shared with the development transcripts). |
For the competition, each clip of this release was processed further before training:
bshall/knn-vc, MIT, © 2023 MediaLab, Stellenbosch University; WavLM-Large features, MIT; prematched HiFi-GAN vocoder; top-k = 4). Each clip was converted toward one adult speaker of the Bangor Miami corpus (GPL-3.0-or-later; BangorTalk, also on MDC cmmfulo4r018bnz07py4q9t09). The matching set of each speaker came from that speaker's own clean utterances in the corpus (at least 1.5 s long, no overlapping speech, training sessions only), 180–480 s per speaker. Each synthetic persona was mapped to one target speaker by a fixed hash of the persona name, so a persona keeps one voice; the field target_voice gives this assignment as opaque IDs. Words and timing were not changed. The same round-trip QC was applied (2,319 clips kept, QC WER 8.85 %).@ID headers) were removed: 86 utterances from 2 speakers. All 37 remaining target voices are adults.The code that rebuilds the training copy from this dataset and the Bangor Miami corpus will be released with the competition solution. The exact training copy can be provided to the competition organisers for verification.
ds3_synthetic_es_en_no_vc/
wav/syn_{conv}_{turn}_{persona}_16k.wav # 2,233 files, 16 kHz mono
manifest.jsonl # one line per utterance
LICENSE (GPL-3.0-or-later) · NOTICE · README.md (this datasheet) · SHA256SUMS
Manifest fields: audio (relative path) · transcript (orthographic text) · duration (s) · conversation · turn_role (A/B) · tts_persona · target_voice (opaque ID T01–T37, the same IDs as in the competition training copy).
Standard orthography with case, accents and punctuation as generated (e.g. ¿…?, ¡…!). Spanish and English words are written in their own spelling. There are no language tags inside the text. Disfluencies and fillers appear as words when the LLM produced them (≈ 8 per 1,000 words). Numbers are mostly written as words (0.9 % of utterances contain digits).
syn_0012_01_meksiko0a_16k.wav · 4.2 s — I think the el precio es but it vale la pena maybe you agreesyn_0284_06_meksiko2b_16k.wav · 5.2 s — ¿Tú traerás also cheese? ¿me lo dices? anyway now heresyn_0001_03_tengah1b_16k.wav · 5.2 s — I think you should but no te olvides to update your résumé porque los recruiters love details.syn_0238_00_meksiko0a_16k.wav · 5.8 s — ¿Sabes que perdí las llaves mami y ahora estoy atrasada porque el metro está lento y no sé si me puedes ayudar para ah tú me...?syn_0092_89_karibia7b_16k.wav · 6.7 s — Sí pienso que el plan es good pero me preocupa la tasa de interés, y el hecho de que no incluyan garantíasyn_0331_03_selatan2b_16k.wav · 12.3 s — I can check the circuit board, but primero I need the manual y si no lo tienes we can also call the power company y si no ¿puedes pasarme el número del servicio de energía para llamar?syn_0128_07_tengah7b_16k.wav · 10.4 s — Oye ¿tienes idea si el doctor está en el consultorio ahora but the waiting room es grande y I think we might need to wait y qué te parece si nos encontramos antes de hoy?syn_0334_00_tengah2a_16k.wav · 15.2 s — Hey mijo I saw the shift chart and I think we got schedule today so mi turno es de ocho a diez pero que me preocupa si el día viernes es de doce a diez, I think this might cause a problem for me y si me dan el día libre, podría viajar tarde para la casa de mi abuela de nuevo.Deuchar, M., Davies, P., Herring, J. R., Parafita Couto, M. C., & Carter, D. (2014). Building bilingual corpora. In E. M. Thomas & I. Mennen (Eds.), Advances in the Study of Bilingualism (pp. 93–110). Multilingual Matters.
Compensated · similar languages
High-specularity optical edge-case dataset (425 assets) for benchmarking computer vision and autonomous navigation against severe specular reflections.
Versioned software-engineering corpus with code, configuration, workflows, and AI-tooling artifacts across frontend, backend, cloud, and infrastructure.