Related Experiment Video
Updated: Aug 26, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Acoustic Characterization and Feature Resistance of Mechanical Versus Memory Voice Disguises: A Pilot Study
Ricardo Buoso Neto1, Julio Cesar Cavalcanti2, Plinio A Barbosa1
1Institute of Language Studies, University of Campinas, Campinas, SP, Brazil.
Objectives:
This study explores the impact of two distinct disguise strategies (ie, mechanical obstruction and disguising from memory) on acoustic-phonetic features and speaker discrimination accuracy. We hypothesized that acoustic parameters most effective at characterizing a disguise would be the least robust for speaker discrimination.
Methods:
Speech samples were collected from nine speakers (five female, four male) under three conditions: baseline, mechanical (eg, bite-block), and memory (eg, pitch shift). We performed a univariate analysis to identify high-performing acoustic features and a multivariate analysis.
Results:
Using Euclidean nearest-neighbor distances, speaker discrimination was higher in the Mechanical condition (hit-rate 0.778; 95% CI [0.400, 0.972]) than in the Memory condition (0.333; [0.075, 0.701]) in this N=9 sample. In mechanical, the top-ranked univariate DRI features for the combined cohort were fundamental-frequency distribution measures (eg, baseline f0), whereas in memory, the top-ranked features shifted toward dynamic pitch derivatives and spectral/formant measures (eg, negative pitch slopes, spectral tilt, and median F2), with wide uncertainty for several estimates due to the small sample.
Conclusions:
The trends are consistent with a hypothesized trade-off in which discrimination may be more successful when the identity-anchoring features are less perturbed by the main target of the disguise mechanism (source vs filter). In this pilot sample, mechanical disguises appeared to distort articulatory features (formants) while leaving phonatory features (f0) relatively less affected for discrimination, whereas the memory-based disguise appeared to alter prosodic and voice-quality traits more substantially.

