Related Experiment Video
Updated: Sep 11, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
A Comprehensive Polish Medical Speech Dataset for Enhancing Automatic Medical Dictation
Andrzej Czyżewski1, Sebastian Cygert1, Karolina Marciniuk1
1Gdańsk University of Technology, Multimedia Systems Department, Faculty of Electronics, Telecommunications and Informatics,, Gdańsk, Poland.
Abstract:
Pre-trained models have become widely adopted for their strong zero-shot performance, often minimizing the need for task-specific data. However, specialized domains like medical speech recognition still benefit from tailored datasets. We present ADMEDVOICE, a novel Polish medical speech dataset, collected using a high-quality text corpus and diverse recording conditions to reflect real-world scenarios. The dataset includes domain-specific vocabulary such as drug names and illnesses, with nearly 15 hours of audio from 28 speakers, including noisy environments. Additionally, we release two enhanced versions: one anonymized for privacy-sensitive use and another synthetic version created via text-to-speech, totaling over 83 hours and nearly 50,000 samples. Evaluating the Whisper model, we observe a 24.03 WER on our test set. Fine-tuning with human recordings reduces WER to 15.47, and incorporating anonymized and synthetic data further lowers it to 13.91. We open-source the dataset, fine-tuned model, and code on Kaggle to support continued research in medical speech recognition.
More Related Videos
Related Concept Videos
Methods of Documentation II: POMR
Data Reporting and Recording

