Release Date: 7/28/2026
Format: WAV, TXT
Size: 4.61 GB
The StammerTalk dataset was created by StammerTalk (口吃说) community (http://stammertalk.net/), in partnership with AImpower.org. The dataset comprises approximately 43 hours of spontaneous conversation and voice-command reading from 66 Mandarin speakers who stutter. The data collection and annotation were led by the StammerTalk community. In partnership with the StammerTalk community, AImpower.org's analysis and research on this dataset demonstrated its utility in benchmarking and fine-tuning state-of-the-art automatic speech recognition (ASR) models for stuttered speech. See: 1. Jingjin Li, Qisheng Li, Rong Gong, Lezhi Wang, and Shaomei Wu. 2025. Our Collective Voices: The Social and Technical Values of a Grassroots Chinese Stuttered Speech Dataset. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT '25). Association for Computing Machinery, New York, NY, USA, 2768–2783. https://doi.org/10.1145/3715275.3732179 2. Rong Gong, Hongfei Xue, Lezhi Wang et al. 2024. AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection. Interspeech 2024. https://arxiv.org/abs/2406.07256
Share
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is intended to support academic and scientific research on speech technologies that are inclusive of people who stutter. Users are asked to use the dataset in ways that advance accessibility, fairness, inclusion, and understanding of stuttered speech and other speech disabilities. Users are asked to publicly disclose, where appropriate, their affiliation, intended use cases, and resulting publications, products, or services developed using the dataset. Users are asked to implement reasonable technical, administrative, and organizational safeguards to protect contributor privacy and reduce the risk of misuse, unauthorized disclosure, or re-identification. Users are asked to notify AImpower.org (contact@aimpower.org) and StammerTalk (ontact@globalchinesestuttering.org) of significant downstream uses, publications, commercial deployments, or other publicly released works developed using the dataset.
Forbidden Usage
Users are asked not to attempt to identify, contact, or otherwise determine the identity of contributors represented in the dataset. Users are asked not to use the dataset to develop, evaluate, or deploy systems that stigmatize, discriminate against, surveil, disadvantage, or otherwise harm people who stutter or other speech-disabled individuals. Users are asked not to use the dataset in ways that are materially inconsistent with the purposes communicated to contributors at the time of data collection. Users are asked not to create synthetic content, voices, avatars, or other outputs intended to impersonate, misrepresent, exploit, or falsely attribute speech or characteristics to contributors. Users are asked not to use the dataset to create commercial models, systems or products. Nothing in the dataset license should be understood to authorize conduct that would violate applicable privacy, publicity, data protection, anti-discrimination, consumer protection, or other laws.
Ethical Review
All contributors provide informed consent, with clear explanations of how their data may be used and shared. We have worked with counsel to develop robust contributor documentation. Personal and identifiable information is removed during the uplift process, and contributors are de-identified to protect privacy. Data governance decisions are made by the StammerTalk community, emphasizing respect, autonomy, and harm minimization.
Intended Use
- Improving speech technology systems for disfluent speakers; - Supporting clinical research and training on speech language pathology; - Advocating for broader public awareness and acceptance of speech diversity; - Supporting academic and scientific research on speech science and technologies; - Developing commercial products to benefit the stuttering community.
Speech data collection was conducted by two StammerTalk volunteers, who also stutter, with participants over videoconferencing platforms. The recorded speech contains both unscripted conversations between the volunteer and the participant, and the dictation of a list of 200 voice commands by the participant. Total 70 adults who stutter (AWS) participated in the recording with two StammerTalk volunteers, resulting in a dataset of 48.8 hours speech from 72 AWS. The speech data from 64 participant and two volunteers were consented to be included in this dataset.
The recorded speech was transcribed semantically and verbatim, with five distinct stuttering event annotations embedded in markups. Obtaining verbatim transcription that includes word repetitions (e.g. “My, my, my name”) and interjections (e.g. “hmm”) was a deliberate choice made by the StammerTalk community, to allow disfluencies rather than automatically erased by ASR models. The annotation was performed by professional speech data annotators, and reviewed by a StammerTalk volunteer.
More details on data collection process, as well as its community impact, can be found in our CSCW '24 and Interspeech 2024 papers:
Qisheng Li and Shaomei Wu. 2024. "I Want to Publicize My Stutter": Community-led Collection and Curation of Chinese Stuttered Speech Data. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 475 (November 2024), 27 pages. https://doi.org/10.1145/3687014
Rong Gong, Hongfei Xue, Lezhi Wang et al. 2024. AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection. Interspeech 2024
The speech was manually annotated by professional speech annotation service providers, under the supervision of the StammerTalk community.
Both verbatim and semantic transcription were created, with embedded markups for five types of stutters specified by the annotation guidelines, including:
[]: Word-level repetition. Repeated words or phrases.
/r: sound repetition. Repeated sounds, such as a consonant or vowel, that do not constitute an entire word.
/b: blocks. Prolonged blocks or unnatural silence.
/p: prolongation. Prolonged phonemes.
/i: interjection. Excessive utterances like ’嗯’ (hmm), ’啊’ (ah), or ’呃’ (um). Notably, natural sounding interjections that do not disrupt the speech flow are excluded from this category.
Example:
Annotation: 我叫[我叫/p]小/b明,我[我我]住/p/b在呃/i北/r京
Interpretation: I am I am (multi-words repetitions and prolongation of “am”) Xiao (block) Ming, I I I (single word repetition) live (prolongation and block) in um (interjection) Bei (“b” sound repetition) Jing.
More details on data annotation process can be found in our CHI '24 paper and Interspeech 2024 papers:
Jingjin Li, Qisheng Li, Rong Gong, Lezhi Wang, and Shaomei Wu. 2025. Our Collective Voices: The Social and Technical Values of a Grassroots Chinese Stuttered Speech Dataset. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT '25). Association for Computing Machinery, New York, NY, USA, 2768–2783. https://doi.org/10.1145/3715275.3732179
Rong Gong, Hongfei Xue, Lezhi Wang et al. 2024. AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection. Interspeech 2024
Scale: despite broad community participation, this dataset consists of less than 50 hours of speech from 45 speakers – a relatively small scale comparing to speech data used to train dominant ASR models. While our experiment showed significant performance gain by fine-tuning mainstream ASR models using a relatively small amount of data (30 hours), more data might be needed for develop ASR models that are inclusive by default.
Disfluency severity: despite targeted outreach, less than 10% of the speakers in our dataset exhibited severe stuttering. While this matches with the natural distribution within the stuttering community, we see a need to over-sample speakers with more severe disfluency to extend the data heterogeneity for speech AI systems.
Geographical and language representation: our Mandarin stuttered speech dataset predominantly represents the Chinese-speaking stuttering community, with the majority of data contributors residing in mainland China and speaking Mandarin Chinese. Other Chinese languages and dialects were not captured in this dataset.