Release Date: 9/14/2026
Format: FLAC, JSON, TSV
Size: 8.10 GB
140 hours of spontaneous, unscripted Brazilian Portuguese speech from 512 speakers, recorded during real job interviews conducted by an AI interviewer. This is not read speech and not acted dialogue. Every speaker was a genuine candidate answering an interviewer's questions about their own work, with a real job on the line. The corpus therefore carries what prepared-script corpora do not: hesitation and self-correction, filler ("nΓ©", "entΓ£o", "tipo"), variable pacing, spontaneous reformulation, and consumer-grade acoustics. 26% of the material was recorded on phones, over domestic Brazilian broadband, in whatever room the speaker had available. 7,043 utterance-level clips, one per interview answer, median 56 seconds. Sentence-segmented transcripts with start and end timestamps ship with every clip, alongside speaker region and city, device class, speech rate, employer industry and job function. The dataset is scrubbed for personally identifying information before release. Speaker names, third-party person names, source employer names and spoken contact details are removed from both the audio and the transcripts. Technology and product names, cities and regions are deliberately retained - they identify nobody, and they are much of what makes the corpus useful. Recorded 1 March - 24 August 2026 across 21 Brazilian states and 112 cities. Audio is 16 kHz mono FLAC, lossless.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($560.00) platform fee paid to MDC.
Restrictions/Special Constraints
Access is granted by request only. This dataset is scrubbed for personally identifying information: speaker names, third-party person names, source employer names and spoken contact details are removed from both the audio and the transcripts. It is not, and cannot be, anonymous - a voice is itself a biometric identifier, and the protection that matters is the prohibition below, not the scrub. Any attempt to identify, re-identify or contact the speakers featured in this dataset is strictly forbidden, including cross-referencing against other datasets and voice-print matching. The dataset may not be redistributed, resold or sublicensed, in whole or in part. Speakers consented to AI training use and may withdraw that consent; on notice, withdrawn material must be deleted from the licensee's copies and excluded from further training.
Forbidden Usage
You agree not to attempt to determine the identity of, re-identify, or contact any speaker in this dataset. Any attempt to clone the voice of, or train models that imitate, an identifiable speaker in this dataset is forbidden. It is forbidden to use this dataset to train or evaluate biometric identification, face recognition or speaker-verification systems. It is forbidden to use this dataset to infer personal attributes about individual speakers, or in any way that could compromise their privacy or employment. Redistribution, resale or sublicensing of the dataset, in whole or in part, is forbidden.
Ethical Review
Every recording in this dataset carries an explicit, affirmative consent flag for use of the material in AI training. Consent is recorded against each submission with a timestamp and the point at which it was given. It is a separate decision from applying for the role: consenting is not a condition of being interviewed or of being considered for the job, and declining does not affect the application. No recording without that flag is included. Nothing here is scraped, and nothing is repurposed from public video. All material was captured in-product, on the speaker's own device, with the speaker aware they were being recorded. Speakers were not paid for the recording; they were candidates interviewing for a role. Consent is revocable at any time and withdrawal is honoured retroactively: withdrawn material is removed from the corpus and excluded from future deliveries, and licensees are notified so they can delete it. Limits of this process, stated plainly: the dataset has not been reviewed by an external institutional ethics board β the controls above are the operator's own governance. The consent record is a flag, a source and a timestamp; the exact wording of the consent notice in force is not versioned per submission, but the copy in force across the collection window (1 March - 24 August 2026) can be supplied on request.
Intended Use
Brazilian Portuguese ASR training and evaluation, particularly for systems that must handle spontaneous, unscripted speech rather than read or acted material. Also suitable for: disfluency and hesitation modelling; robustness testing against consumer-grade capture (phone microphones, domestic broadband, untreated rooms); speech-rate and prosody research on unprepared speech; forced alignment and segmentation research; and conversational or multimodal model training where genuine spontaneous Brazilian Portuguese is required.
140 hours of spontaneous, unscripted Brazilian Portuguese speech from 512 speakers, captured during real job interviews conducted by an AI interviewer between 1 March and 24 August 2026. 7,043 utterance-level clips, one per interview answer, each with a timestamped machine transcript.
All figures on this page are a frozen snapshot taken on 25 August 2026. The underlying corpus continues to grow; this dataset is a fixed cut so that the listing, the archive and the manifest agree with one another. Later cuts will be published as new versions rather than as silent edits to these numbers.
Speakers were genuine candidates answering an interviewer's questions about their own work, with a real job at stake. That is what separates this material from read-speech and acted-dialogue corpora: it carries hesitation, self-correction, genuine filler (ne, entao, tipo), variable pacing, mid-sentence reformulation, and the acoustics of whatever room the speaker happened to be in.
Capture was browser-based, over the public internet, on the speaker's own device. No studio, no headset requirement, no noise-floor guarantee. Real-world conditions are the point of the dataset, not a defect in it.
| Metric | Value |
|---|---|
| Unique speakers | 512 |
| Clips | 7,043 (about 14 per speaker) |
| Total audio | 140.0 hours |
| Per speaker | median 15.0 min, mean 16.4 min, p90 27.3 min, max 69.5 min |
No gender or age labels. These attributes are not collected and are not inferred.
Share of total hours by device class:
| Device | Share of hours |
|---|---|
| Windows desktop | 65.0% |
| Android phone | 13.8% |
| iPhone | 11.8% |
| Mac | 7.1% |
| Linux / ChromeOS | 2.3% |
Roughly 26% of the material is phone-recorded, over domestic Brazilian broadband.
| Property | As captured |
|---|---|
| Container | MP4 97%, WebM 3% |
| Audio codec | AAC 73%, Opus 27% |
| Sample rate | 44.1 kHz 50%, 48 kHz 50% |
| Channels | 2-channel 90%, mono 10% |
| Audio bitrate | approx. 127 kbps mean |
Delivered as 16 kHz mono FLAC, lossless, approximately 8 GB for the full cut. Synchronised video is retained by the owner and is not part of this listing, so no speaker's face is included in the delivered artefact.
| Metric | Value |
|---|---|
| Clips | 7,043 (one per interview answer) |
| Duration | median 56.4 s, mean 70.9 s |
| Duration spread | p25 34.8 s, p75 89.3 s, p90 130.1 s, max 25.3 min |
| Shortest retained | 4.9 s (clips under 3 s are excluded) |
| Speech rate | median 119 wpm, p10 79, p90 157 |
| Transcribed words | approx. 1.0 million |
Sentence-segmented with start and end timestamps, supplied as one JSON per clip plus a flat TSV.
Machine-generated and not verified by a human listener. These are the output of the interview platform's live speech-to-text, captured during the interview itself, rather than a batch re-transcription. No WER figure is quoted for this release because none has been measured on this cut.
Sentences muted by the PII scrub read [REDACTED] and their audio carries room tone, so the transcript and the audio always agree.
21 states and 112 cities. The distribution is concentrated, and that concentration is disclosed rather than smoothed over.
| State | Hours | Share |
|---|---|---|
| Parana | 66.8 | 47.0% |
| Sao Paulo | 43.2 | 30.4% |
| Rio de Janeiro | 5.4 | 3.8% |
| Amazonas | 4.8 | 3.4% |
| Minas Gerais | 4.2 | 2.9% |
| Santa Catarina | 3.7 | 2.6% |
| Rio Grande do Sul | 2.8 | 2.0% |
| Bahia | 2.1 | 1.4% |
| 13 further states | 11.2 | 8.9% |
Top cities: Curitiba 49.0 h (34.5%), Sao Paulo 25.0 h (17.6%), Manaus 4.8 h, Sao Jose dos Pinhais 3.7 h, Santo Andre 3.1 h, Rio de Janeiro 3.0 h.
Dialect is inferred from country of recording. The audio has not been dialect-annotated by a listener. 100% of the material was recorded in Brazil.
| Domain (employer industry) | Share |
|---|---|
| IT and software | 42.9% |
| Professional services | 23.9% |
| Unclassified | 16.2% |
| Media and advertising | 9.9% |
| Finance | 2.5% |
| Manufacturing, retail, hospitality, supply chain, healthcare, other | 4.6% |
1,173 distinct interview questions across 34 interview sets. Roles range from software and infrastructure internships through procurement analysts, security technicians, marketing and administrative coordinators, to talent-acquisition and sales positions.
Every clip ships with: speaker region and city, device class, clip duration, redacted seconds, speech rate (wpm), employer industry, job function, an opaque employer-group key, and the interview question the answer responds to (with employer names removed).
Candidate names, email addresses, CV text, employer names, interview-set titles and source media URLs are held by the dataset owner and are not shipped in any form. The source video is not shipped.
Collected on Flowmingo, an AI interview platform used by employers to screen candidates. Every speaker gave explicit, affirmative consent for AI-training use as a decision separate from, and not a condition of, their job application. Nothing is scraped and nothing is repurposed from public video. Consent is revocable and withdrawal is honoured retroactively.
Source employers are not named. The cut draws on 8 employers; the largest single employer accounts for 67% of the hours and the two largest together account for 89%, which is why Parana and Curitiba dominate the geography above.
Stated plainly, because a buyer will find them anyway:
Concentrated, not nationally representative. Two employers account for 89% of the hours; Parana and Sao Paulo for 77%. The North and Northeast are thinly covered.
Scrubbed, but not anonymous. Speaker names, third-party names, source employer names and spoken contact details are removed from the audio and the transcripts. A voice is nonetheless a biometric identifier, so this dataset is not anonymised data and must not be treated as such. The re-identification prohibition in the licence is the operative control.
Machine transcripts, live-STT provenance, no measured WER. Not human-verified, and no error rate is claimed for this cut. Priced accordingly.
Topic range is narrow by design. Workplace and professional discourse - career history, technical explanation, situational judgement. Not conversational chat, read speech, or broad-domain content.
Single-channel speaker audio. The interviewer's audio is not retained.
Consent is a flag, not a versioned string. Each submission records a consent flag, its source and a timestamp. The exact notice wording is not versioned per submission; the copy in force across the collection window can be supplied on request.
No dialect annotation and no speaker demographics. See above.
1.41% of the audio is redacted room tone. Muting is at sentence granularity, so a sentence containing a name is removed in full. redacted_sec in metadata.tsv gives the exact figure per clip. 83 clips carrying no transcript were excluded from the release entirely, since they could be neither PII-scanned nor used for ASR training.