Release Date: 9/25/2026
Format: PARQUET, CSV, JSON, SQL, MD
Size: 922.38 MB
The English 10 Hours Medical Speech Evaluation Dataset is a speech dataset designed for evaluating automatic speech recognition (ASR) and other speech-processing systems in medical and healthcare-related contexts. It contains approximately 10 hours of English-language medical speech representing terminology, phrases, and conversational patterns commonly encountered in healthcare settings. The dataset is intended to support the evaluation of how accurately speech recognition systems handle medical vocabulary, clinical expressions, and healthcare-related spoken language. It can be used as an evaluation resource for benchmarking model performance, identifying recognition errors, and assessing robustness in medical speech scenarios. This dataset may also support research and development in medical speech technology, including speech recognition, transcription, and language-processing applications. Researchers should review the accompanying metadata, licensing terms, and usage restrictions before using the dataset.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($50.00) platform fee paid to MDC.
Licensing
MirasAI Custom Data License This dataset is licensed by MirasAI LLC, not public domain or open-license. By using it, you agree: Grant: Non-exclusive, non-transferable license to use this dataset for internal/commercial purposes, including training and fine-tuning AI/ML models and creating derived outputs. Restrictions: No reselling, redistributing, or sublicensing this dataset. No rebuilding a substitute dataset from it. No re-identifying anonymized speakers. Keep attribution notices intact. Ownership: MirasAI retains all rights to the dataset. You own any models/outputs you create using it, provided they don't expose the underlying data. Attribution: Credit "MirasAI LLC" in any resulting publication or product. Warranty: Provided "as is," no warranties of any kind. Governing Law: State of Indiana, USA. Contact: [email protected]
Restrictions/Special Constraints
Use of this dataset is subject to its applicable license and terms of access. Users must not redistribute, resell, or share the dataset without appropriate authorization.
Forbidden Usage
The dataset must not be used for unauthorized commercial purposes, surveillance, or any application that could cause harm to individuals. It must not be used to make medical diagnoses or clinical decisions without appropriate validation and professional oversight.
Intended Use
The dataset is intended for evaluating and benchmarking English medical speech recognition and transcription systems. It may support research on ASR performance, medical terminology recognition, and healthcare-related speech processing.
The English 10 Hours Medical Speech Evaluation Dataset is a held-out evaluation set of approximately 10 hours of English speech in the medical domain. It is intended for benchmarking automatic speech recognition (ASR) and related speech models on clinical and medical vocabulary.
The dataset is built for evaluation only. It comes with exclusion lists so that training data can be kept separate from the evaluation audio and its medical terms. That way, reported scores measure how well a model generalises rather than how much it has memorised.
Language: English
Domain: Medical / Healthcare
English_10_Hours_Medical_Speech_Evaluation_Dataset/
├── data/
│ ├── evaluation-00000-of-00009.parquet
│ ├── evaluation-00001-of-00009.parquet
│ ├── evaluation-00002-of-00009.parquet
│ ├── evaluation-00003-of-00009.parquet
│ ├── evaluation-00004-of-00009.parquet
│ ├── evaluation-00005-of-00009.parquet
│ ├── evaluation-00006-of-00009.parquet
│ ├── evaluation-00007-of-00009.parquet
│ └── evaluation-00008-of-00009.parquet
├── CHECKSUMS.sha256
├── DATASET_CARD.md
├── excluded_terms.csv
├── LICENSE.md
├── manifest.csv
├── PROVENANCE.md
├── README.md
├── selection_method.sql
├── selection_summary.json
├── shard_inventory.csv
├── training_exclusions_exact_audio.csv
└── training_exclusions_term_disjoint.csv
| Field | Value |
|---|---|
| Dataset Name | English 10 Hours Medical Speech Evaluation Dataset |
| Language | English |
| Domain | Medical |
| Split | Evaluation |
| Total Duration | ~10 hours |
| Number of Shards | 9 |
| File Format | PARQUET, CSV, JSON, SQL, MD |