Release Date: 7/23/2026
Format: Parquet
Size: 5.69 GB
Share
This dataset contains 11,563 audio-text pairs of spoken English medical content, covering clinical abbreviations, diagnoses, and terminology across specialties including cardiology, gastroenterology, neurology, obstetrics, and emergency medicine. Each entry pairs an audio recording with its corresponding transcript, structured as educational sentences explaining medical terms and their clinical context (e.g. abbreviation expansions, diagnostic criteria, treatment approaches). The dataset is intended for training and evaluating medical-domain automatic speech recognition (ASR) and speech-to-text systems.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is licensed under CC-BY-NC-SA-4.0 and is intended for non-commercial research use only, including training and evaluation of medical-domain ASR/speech-to-text models. Attribution to Proxima AI is required, and any redistributed derivative must be released under the same license. This dataset provides educational medical content and should not be used as a substitute for verified clinical or diagnostic information.
Forbidden Usage
Commercial use of this dataset, or of any model trained on it, without prior written permission from Proxima AI. Redistribution without retaining attribution to Proxima AI. Use of this dataset's medical content as authoritative clinical guidance without verification by a licensed medical professional.
Total samples: 11,563
Fields: path (audio), sentence (transcript text), client_id (speaker identifier), language
Total size: 7.04 GB
Medical education content covering clinical abbreviations and terminology across specialties including cardiology, gastroenterology, neurology, obstetrics, orthopedics, and emergency medicine. Each sentence explains a medical term/abbreviation with its clinical context and significance.
Used in training mahwizzzz/medwhishper, a 0.2B-parameter ASR model.
Released under CC-BY-NC-SA-4.0. Non-commercial use only, attribution to Proxima AI required, derivatives must be shared under the same license.