Related Experiment Video
Updated: Aug 18, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
The development of an automatic speech recognition model using interview data from long-term care for older adults
Coen Hacking1,2, Hilde Verbeek1,2, Jan P H Hamers1,2
1Department of Health Services Research, CAPHRI Care and Public Health Research Institute, Faculty of Health Medicine and Life Sciences, Maastricht University, Maastricht, The Netherlands.
Improving automatic speech recognition (ASR) for older adults in long-term care is crucial. Fine-tuning ASR models with interview data significantly reduces errors and speeds up transcription, enhancing data collection for care quality research.
Area of Science:
- Gerontology
- Speech Technology
- Health Informatics
Background:
- Interviews in long-term care (LTC) are vital for understanding older adults' perspectives.
- Manual transcription of these interviews is time-consuming and labor-intensive.
- Existing automatic speech recognition (ASR) systems often perform poorly with specific demographics, including older adults and accented speech.
Purpose of the Study:
- To develop an effective ASR model tailored for interview data from older adults in LTC settings.
- To demonstrate the efficacy of using specific demographic data to improve ASR performance.
- To reduce the word error rate (WER) and time required for transcribing qualitative data in LTC research.
Main Methods:
- An initial ASR model was created using the Mozilla Common Voice dataset.
- The model was fine-tuned using 34 hours of audio and transcript data from interviews with LTC residents, families, and staff.
- Continuous processing and refinement of interview data were employed to minimize the word error rate (WER).
Main Results:
- The initial ASR model had a high WER of 48.3% on interview data due to background noise and mispronunciations.
- Fine-tuning with interview data reduced the average WER to 24.3%, with a median WER of 22.1% on test data.
- The improved ASR model was over six times faster than manual transcription, with residents' speech showing the highest WER at 22.7%.
Conclusions:
- Fine-tuning ASR models with specific interview data significantly decreases the word error rate, proving the method's effectiveness.
- The developed ASR system offers a substantial time-saving solution for transcribing qualitative data in long-term care research.
- Local transcription using the improved ASR model can enhance participant privacy by minimizing data sharing.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
06:04Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023