indospeech
- Indonesia
- Https://www.Indospeech.com
- Sole Proprietorship
About indospeech
I run spontaneous conversational audio collection in West Java, Indonesia, and I am writing to offer myself as an in-country collection partner rather than a one-off dataset seller.
WHAT I COLLECT
Three matched conditions from the same speaker pool and the same recording setup:
Monolingual Sundanese (su-ID)
Monolingual Indonesian (id-ID)
Natural Sundanese-Indonesian code-switching, unprompted
That third condition is the one most catalogues miss. Code-switching is where production ASR breaks, and it cannot be simulated by recording two separate monolingual corpora. Having all three from the same speakers, region and setup gives you a controlled comparison rather than three unrelated datasets.
HOW IT IS RECORDED
One external lavalier microphone and one separate recording device per speaker
48 kHz, 16-bit WAV masters, mono per channel
2 to 6 speakers per session
Entirely unscripted: no prompts, no topic list, no script
Natural overlap, interruption, laughter
Most Indonesian conversational data on the market is 16 kHz mono with mixed channels and topic-guided dialogue. Separate channels per speaker mean diarization ground truth rather than a model guessing who spoke.
HOW QUALITY IS MEASURED
Every session is analysed before delivery. A script finds passages where only one speaker is active and measures, in decibels, how much quieter every other channel is. The worst case sets the grade, and that figure ships in the metadata so your QA team can verify rather than trust.
Current reference session: channel isolation 17.88 dB, SNR 53.5 dB, clipping 0.0% of samples. Grade B by my own scale, stated openly. Grade A is above 18 dB isolation.
CONSENT
Bilingual written consent signed by every participant before recording, read aloud and explained in Sundanese or Indonesian. Covers commercial AI training, sublicensing, transcription and annotation by third parties, voice synthesis and cloning, and cross-border transfer under Indonesian Law No. 27/2022.
Records are held in two parts. Part A carries the name and signature and never leaves us. Part B carries only an anonymous speaker code, gender and age range, and ships with the data. Your legal team can verify consent exists without receiving personal identifiers.
Datasets
No datasets published yet.