Related Experiment Video
Updated: Sep 28, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Development and evaluation of a voice dataset for voice therapy exercises: quantitative audio analysis and user
Sabrina H Schröder1, Rica Schulze2, Merle Schlender1
1University Clinic for Visceral Surgery, Carl von Ossietzky University Oldenburg, Oldenburg, Germany.
Purpose:
This study aims to develop and validate a high-quality voice dataset for dysphonia therapy using a standardized, video-guided recording protocol comprising a range of therapeutic voice exercises. The dataset is designed to support the development of digital health applications that enable home-based voice training with real-time biofeedback, thereby addressing current challenges in speech-language pathology, including limited access to therapy and the need for scalable, patient-centered interventions.
Methods:
A voice dataset was collected from adult participants performing voice therapy exercises in German, recorded via smartphone and in-ear microphones under standardized conditions. In addition to acoustic data collection, participants evaluated the recording procedure and video-guided exercise instructions. Acoustic analyses of pitch (Hz), intensity (dB), and Cepstral Peak Prominence (CPP) were conducted using Praat, with data grouped by age and sex. Correlation analyses were performed to assess the consistency of recordings across participants and exercises.
Results:
The resulting dataset comprises recordings from 175 adults, totaling approximately 6.6 GB of audio data, with a subset made available as open access. Participants reported high overall acceptance of the recording procedure and instructional videos, while also identifying areas for improvement, particularly regarding pacing and clarity. Acoustic and correlation analyses demonstrated exercise-specific variations in CPP, pitch and intensity, as well as systematic differences across age and gender groups, indicating both variability and structured consistency within the dataset.
Conclusion:
The dataset represents a novel resource for digital voice therapy, particularly due to its focus on structured therapeutic exercises in German. It provides a robust foundation for the development and evaluation of digital health applications aimed at enhancing self-managed voice training and supporting therapy processes beyond clinical settings.