Release Date: 7/28/2026
Format: WAV, TXT
Size: 9.75 GB
The dataset is organized into two folders: Audio and Annotation. Each folder contains 44 sub-folders, named by participant ID (e.g. P01, P11). The annotations are timestamped, containing both the semantic and verbatim transcriptions of what was said, as well as five stuttering event annotations embedded as markups. The verbatim transcripts include stuttered repetitions (e.g. “My, my, my name”, “Th-th-thank you”) and interjections (e.g. “hmm”, “uh”). The annotation guidelines were co-designed with people who stutter and speech language pathologists (SLPs) working closely with people who stutter. The transcription and annotation was performed by annotators with SLP background, and reviewed by SLP specialized in stuttering to ensure rigor and consistency. The dataset was developed through community-led governance models, with contributors actively involved in decisions around data use, access, and representation. The data governance model was informed by stuttering researchers, prominent stuttering advocates, and community organizers, as well as the broad stuttering community in the US, Canada, and China. Read our team's AIES '25 paper for more details on the governance principles and process: - Li, J., Liu, P., Lietz, R., Tang, N., Su, N. M., & Wu, S. (2025). Govern with, Not For: Understanding the Stuttering Community’s Preferences and Goals for Speech AI Data Governance in the US and China. Proceedings of the AAAI ACM Conference on AI, Ethics, and Society, 8(2), 1548–1560. https://doi.org/10.1609/aies.v8i2.36654 51 speakers who stutter participated in the data collection. Participant P03 later withdrew their conversational data from the final dataset. We also withhold all the speech data from Participants P06, P07, P09, P35, P37, P41 from this dataset to be used as test data for future contests/competitions.
Share
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
This dataset is intended to support research, evaluation, and development of speech technologies that are inclusive of people who stutter. Users are asked to use the dataset in ways that advance accessibility, fairness, inclusion, and understanding of stuttered speech and other speech disabilities. Users are asked to publicly disclose, where appropriate, their affiliation, intended use cases, and resulting publications, products, or services developed using the dataset. Users are asked to implement reasonable technical, administrative, and organizational safeguards to protect contributor privacy and reduce the risk of misuse, unauthorized disclosure, or re-identification. Users are asked to notify AImpower.org of significant downstream uses, publications, commercial deployments, or other publicly released works developed using the dataset.
Forbidden Usage
Users are asked not to attempt to identify, contact, or otherwise determine the identity of contributors represented in the dataset. Users are asked not to use the dataset to develop, evaluate, or deploy systems that stigmatize, discriminate against, surveil, disadvantage, or otherwise harm people who stutter or other speech-disabled individuals. Users are asked not to use the dataset in ways that are materially inconsistent with the purposes communicated to contributors at the time of data collection. Users are asked not to create synthetic content, voices, avatars, or other outputs intended to impersonate, misrepresent, exploit, or falsely attribute speech or characteristics to contributors. Nothing in the dataset license should be understood to authorize conduct that would violate applicable privacy, publicity, data protection, anti-discrimination, consumer protection, or other laws.
Ethical Review
This study followed established ethical principles for human-subjects research, including informed consent, voluntary participation, privacy protection, and data security. Before participating, participants received and signed a consent form describing the study purpose, procedures, potential risks and benefits, compensation, and data-use practices. Participants were informed that their speech recordings would be used to support the development of more inclusive speech technologies for people who stutter. Participation was voluntary, and participants could withdraw at any time without penalty.
Intended Use
Intended and example use cases includes, but not limited to: - Improving speech technology systems for disfluent speakers; - Supporting clinical research and training on speech language pathology; - Advocating for broader public awareness and acceptance of speech diversity; - Supporting academic and scientific research on speech science and technologies; - Developing commercial products to benefit the stuttering community.
Participants were recruited from the stuttering community. Eligible participants were adults who identified as people who stutter and were proficient English speakers. Prior to participation, all participants reviewed and signed an informed consent form. Data collection was conducted remotely via Zoom. Each session lasted approximately one hour, and participants completed two speech tasks: (1) a 40-min spontaneous conversation with another person who stutters or a stuttering ally to capture natural stuttered speech in conversational settings, and (2) a 20-min voice-command recitation task in which participants read a set of predefined commands.
Audio recordings from both tasks were collected. Participants received $50 compensation upon completion of the recording session.
To protect participant privacy, recordings were de-identified by removing personally identifiable information and assigning participant ID codes. The de-identified data were securely stored by AImpower.org and shared only with approved research and industry partners under data-use agreements and secure data-sharing procedures.
All audios are reviewed and annotated by trained researchers with extensive expertise in speech-language pathology and stuttering. Verbatim transcripts were created and preserved all stuttering events and disfluencies, using a stuttering annotation scheme developed in collaboration with researchers and professionals with expertise in stuttering. The scheme captured five common types of stuttering events: blocks, prolongations, sound repetitions, word/phrase repetitions, and interjections.
Blocks (/b) were marked immediately before the sound or word where a blocking event occurred (e.g., “My /bname”).
Prolongations (/p) were marked immediately after the elongated sound or syllable (e.g., “M/pommy”).
Sound repetitions ([]/s) were used when a sound or syllable was repeated multiple times. The repeated sounds were enclosed in brackets, with repetitions separated by hyphens, followed by the /s marker (e.g., “[pr-pr-pr-]/sprepare”).
Word or phrase repetitions ([]/r) were used when complete words or phrases were repeated. The repeated words were enclosed in brackets and followed by the /r marker (e.g., “[my my]/r my name”).
Interjections ([]/i) were used to annotate filler words or stuttering-related fillers, such as “uh” or “um,” by enclosing them in brackets followed by the /i marker (e.g., “I [uh]/i work”).
More details on data annotation guidelines and process can be found in our CHI 2026 paper:
Xinru Tang, Jingjin Li, and Shaomei Wu. 2026. Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who Stutter. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26). Association for Computing Machinery, New York, NY, USA, Article 284, 1–22. https://doi.org/10.1145/3772318.3790405
Scale: despite broad community participation, this dataset consists of less than 50 hours of speech from 45 speakers – a relatively small scale comparing to speech data used to train dominant ASR models. While our experiment showed significant performance gain by fine-tuning mainstream ASR models using a relatively small amount of data (30 hours), more data might be needed for develop ASR models that are inclusive by default.
Disfluency severity: despite targeted outreach, less than 10% of the speakers in our dataset exhibited severe stuttering. While this matches with the natural distribution within the stuttering community, we see a need to over-sample speakers with more severe disfluency to extend the data heterogeneity for speech AI systems.
Geographical and language representation: our English stuttered speech dataset predominantly represents the English-speaking stuttering community, with the majority of data contributors residing in the United States and speaking with a mainstream English accent. Other English accents or linguistic styles were needed to fully represent the stuttering community and their speech.