Related Experiment Video
Updated: Aug 4, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Mel frequency spectral domain defenses against adversarial attacks on speech recognition systems
Nicholas Mehlman1, Anirudh Sreeram1, Raghuveer Peri1
1Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, California 90089, USA nmehlman@usc.edu, asreeram@usc.edu, rperi@usc.edu, shri@usc.edu.
Abstract:
Automatic speech recognition (ASR) systems are vulnerable to adversarial attacks due to their reliance on machine learning models. Many of the defenses explored for defending ASR systems simply adapt defense approaches developed for the image domain. This paper explores speech-specific defenses in the feature domain and introduces a defense method called mel domain noise flooding (MDNF). MDNF injects additive noise to the mel spectrogram speech representation prior to re-synthesizing the audio signal input to ASR. The defense is evaluated against strong white-box threat models and shows competitive robustness.
Related Concept Videos
IR Frequency Region: Fingerprint Region
Aliasing
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
Frequency-Domain Interpretation of PD Control
The proportional control gain, combined with the...
¹H NMR: Interpreting Distorted and Overlapping Signals
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)

