Related Experiment Video
Updated: Jul 15, 2026

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)
Published on: February 21, 2011
Restoring Over-Attenuated Sustained Vowel Phonations in Smartphone Calls of COPD Patients Using U-Net and Speech
Hyora Lee1, Hyein Ryu2, Sang Mee Lee3
1Department of Digital Health, SAIHST, Sungkyunkwan University, Seoul, Republic of Korea; Medical AI Research Center, Samsung Medical Center, Seoul, Republic of Korea.
Abstract:
Voice data collected through smartphone-based telemedicine may serve as digital biomarkers, but sustained vowel phonation recorded during smartphone calls is often excessively attenuated, limiting its clinical utility. We aimed to characterize this codec-related over-attenuation and develop a restoration framework for attenuated sustained vowel phonation. We retrospectively analyzed 1553 face-to-face sustained vowel phonation samples from 288 outpatients with and without chronic obstructive pulmonary disease and generated paired training data by synthetically degrading clean recordings with time-varying attenuation patterns resembling real calls. We evaluated a 1D U-Net, U-Net with a Whisper encoder, and U-Net with a Wav2Vec 2.0 encoder using a composite loss combining mean squared error, multi-resolution short-time Fourier transform loss, and scale-invariant signal-to-distortion ratio (SI-SDR) loss. General-purpose speech restoration models performed poorly, whereas all proposed models achieved positive SI-SDR values. For full waveform restoration, the U-Net with Whisper achieved the highest mean SI-SDR (13.677 dB), significantly outperforming the U-Net alone (13.190 dB; P = 0.001). In attenuated regions, both foundation model-based architectures significantly improved masked SI-SDR versus the U-Net (both P = 0.001). Qualitative evaluation on real-call samples showed recovery of previously suppressed segments. These findings suggest that speech foundation model-augmented U-Nets can effectively restore codec-attenuated sustained vowel phonation and support more reliable remote voice biomarker collection.
Related Concept Videos
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids, corniculates, and...
Sleep Apnea
The condition is more prevalent among...
Hyperpnea and Hyperventilation
