Release Date: 9/17/2026
Format: PARQUET
Size: 5.69 GB
This dataset contains 11,563 audio-text pairs of spoken English medical content, covering clinical abbreviations, diagnoses, and terminology across specialties including cardiology, gastroenterology, neurology, obstetrics, and emergency medicine. Each entry pairs an audio recording with its corresponding transcript, structured as educational sentences explaining medical terms and their clinical context (e.g. abbreviation expansions, diagnostic criteria, treatment approaches). The dataset is intended for training and evaluating medical-domain automatic speech recognition (ASR) and speech-to-text systems.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($350.00) platform fee paid to MDC.
Licensing
MirasAI Custom Data License This dataset is licensed by MirasAI LLC, not public domain or open-license. By using it, you agree: Grant: Non-exclusive, non-transferable license to use this dataset for internal/commercial purposes, including training and fine-tuning AI/ML models and creating derived outputs. Restrictions: No reselling, redistributing, or sublicensing this dataset. No rebuilding a substitute dataset from it. No re-identifying anonymized speakers. Keep attribution notices intact. Ownership: MirasAI retains all rights to the dataset. You own any models/outputs you create using it, provided they don't expose the underlying data. Attribution: Credit "MirasAI LLC" in any resulting publication or product. Warranty: Provided "as is," no warranties of any kind. Governing Law: State of Indiana, USA. Contact: [email protected]
Restrictions/Special Constraints
This dataset is licensed under MirasAI Custom Data License and is intended for research use only, including training and evaluation of medical-domain ASR/speech-to-text models. This dataset provides educational medical content and should not be used as a substitute for verified clinical or diagnostic information.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it. Use of this dataset's medical content as authoritative clinical guidance without verification by a licensed medical professional.
Intended Use
This English Medical Speech Dataset is intended for research and development in medical speech recognition, automatic speech recognition (ASR), speech-to-text systems, healthcare-related voice technologies, and natural language processing.
A speech dataset of spoken medical education content, capturing clinical abbreviations and terminology across multiple medical specialties. Each recording is a spoken sentence that explains a medical term or abbreviation along with its clinical context and significance.
| Field | Description |
|---|---|
path | Audio file for the recorded sentence |
sentence | Transcript text of the spoken sentence |
client_id | Speaker identifier |
language | Language of the recording (English) |
| Metric | Value |
|---|---|
| Total samples | 11,563 |
| Total size | 7.04 GB |
| Language | English |
Medical education content covering clinical abbreviations and terminology across the following specialties:
Cardiology
Gastroenterology
Neurology
Obstetrics
Orthopedics
Emergency medicine
Each sentence explains a medical term or abbreviation along with its clinical context and significance, making the dataset suited for domain-specific ASR and medical terminology recognition tasks.
| Field | Value |
|---|---|
| Dataset Name | English Medical Speech Dataset |
| Language | English |
| Domains | Cardiology, Gastroenterology, Neurology, Obstetrics, Orthopedics, Emergency Medicine |
| Number of Samples | 11,563 |
| Total Size | 7.04 GB |
| Fields | path, sentence, client_id, language |