Related Experiment Video
Updated: Jul 20, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Gaussian-Filtered High-Frequency-Feature Trained Optimized BiLSTM Network for Spoofed-Speech Classification
Hiren Mewada1, Jawad F Al-Asad1, Faris A Almalki2
1Electrical Engineering Department, Prince Mohammad bin Fahd University, P.O. Box 1664, Al Khobar 31952, Saudi Arabia.
Protecting voice-controlled devices from speech spoofing is crucial. This study introduces an optimized BiLSTM network using high-frequency inverted Mel-frequency cepstral coefficients (iMFCC) for superior spoof detection, achieving 99.58% accuracy.
Area of Science:
- Speech processing
- Machine learning
- Cybersecurity
Background:
- Voice-controlled devices require robust security against speech spoofing attacks.
- Existing spoof detection methods need improvement due to unknown attack algorithms.
- High-frequency speech features show promise in distinguishing genuine from spoofed speech.
Purpose of the Study:
- To develop an effective spoof speech detection model using high-frequency features.
- To investigate the efficacy of Gaussian filters for extracting inverted Mel-frequency cepstral coefficients (iMFCC).
- To optimize a Bidirectional Long Short-Term Memory (BiLSTM) network using a Bayesian algorithm for improved spoof classification.
Main Methods:
- Extraction of high-frequency iMFCC features using a Gaussian filter.
- Integration of complementary features with iMFCC to enhance discrimination.
- Optimization of a BiLSTM network architecture and hyper-parameters via a Bayesian algorithm.
- Training and evaluation on the ASVSpoof 2017 dataset.
Main Results:
- The optimized BiLSTM model achieved 99.58% validation accuracy with minimal epochs.
- The proposed algorithm attained a 6.58% Equal Error Rate (EER) on the evaluation dataset.
- A relative improvement of 78% was observed compared to a baseline spoof-identification system.
Conclusions:
- High-frequency features, particularly iMFCC extracted with Gaussian filters, are effective for spoof speech detection.
- Optimized deep learning models, like the Bayesian-tuned BiLSTM, significantly enhance spoof classification accuracy.
- The proposed method offers a substantial advancement in securing voice-controlled systems against sophisticated spoofing attacks.
More Related Videos
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Aliasing
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
IR Frequency Region: Fingerprint Region
Sampling Continuous Time Signal
In the...

