Related Experiment Video
Updated: Sep 20, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Validation of FUDAN VOICE: Signal and Acoustic Feature Comparison with Computerized Speech Lab (CSL)
1ENT Institute and Department of , Eye & ENT Hospital, Fudan University, 83 Fenyang Road, Shanghai 200031, China.
Objective:
This study aimed to validate the consistency between audio acquired by a custom WeChat mini-program, FUDAN VOICE (named Voice Acquisition in our earlier study), and Computerized Speech Lab (CSL) recordings at both signal and acoustic feature levels, and to characterize aggregate signal deviations between the two recording pathways, encompassing contributions from hardware, platform-level processing, and environmental noise.
Methods:
A simultaneous recording design was employed, capturing parallel audio from CSL and FUDAN VOICE across a soundproof room and a quiet office environment. Signal-level agreement was assessed using the root mean square error (RMSE), Pearson correlation coefficient (Pearson's rho), and mean coherence. Acoustic features such as fundamental frequency (F0), jitter, shimmer, noise-to-harmonic ratio (NHR), cepstral peak prominence (CPP), and cepstral/spectral index of dysphonia (CSID) were extracted and compared for correlation analysis and intraclass correlation coefficients (ICC).
Results:
Signal-domain analysis revealed favorable agreement between FUDAN VOICE and CSL, with low RMSE and preserved spectral structure below 5 kHz, although high-frequency attenuation (> 5 kHz) was observed. Feature-level verification demonstrated excellent concordance for F0 mean, CPP, and CSID (r > 0.82, ICC > 0.82) across environments, whereas F0 extreme values-particularly F0 min-showed marked degradation in non-soundproof settings. Jitter and NHR remained robust, while shimmer exhibited environmental sensitivity.
Conclusions:
FUDAN VOICE achieves reliable remote acoustic acquisition for F0 mean, CPP, and CSID across soundproof and office environments, though high-frequency attenuation (> 5 kHz) and noisy-setting F0 degradation warrant caution. This study establishes the first validation framework for super-app-based voice acquisition, supporting the standardization of WeChat mini-programs in mobile health.

