Related Experiment Video
Updated: Nov 5, 2025

Concurrent EEG and Functional MRI Recording and Integration Analysis for Dynamic Cortical Activity Imaging
Published on: June 30, 2018
Spectral Flux-Based Convolutional Neural Network Architecture for Speech Source Localization and Its Real-Time
Yiya Hao1, Abdullah Küçük1, Anshuman Ganguly1
1Department of Electrical and Computer Engineering, The University of Texas at Dallas, Richardson, TX 75080, USA.
This study introduces a real-time convolutional neural network (CNN) algorithm for robust speech source localization (SSL) in noisy environments. The novel method achieves high accuracy and low latency, improving audio processing for smart devices.
Area of Science:
- Acoustics and Signal Processing
- Artificial Intelligence and Machine Learning
Background:
- Speech source localization (SSL) is crucial for audio processing but challenging in realistic noisy and reverberant conditions.
- Existing SSL algorithms often struggle with performance degradation in complex acoustic environments.
Purpose of the Study:
- To develop a real-time, robust CNN-based SSL algorithm capable of handling realistic background acoustic conditions.
- To evaluate the algorithm's performance on a prototype platform for practical applications.
Main Methods:
- Utilized a convolutional neural network (CNN) trained with features derived from the imaginary-real coefficients of the short-time Fourier transform (STFT) and Spectral Flux (SF).
- Employed delay-and-sum (DAS) beamforming as part of the input feature extraction process.
- Trained the CNN model using diverse noisy speech recordings and tested on unseen acoustic environments.
Main Results:
- The proposed CNN-SSL algorithm demonstrated significant improvements over five previously published methods under various noisy conditions.
- Achieved high accuracy (89.68% at 5dB SNR under Babble noise) with low latency (21 ms per frame).
- Successfully implemented and tested for real-time operation on a Raspberry Pi prototype.
Conclusions:
- The integration of Spectral Flux (SF) with beamforming enhances the CNN's ability to learn temporal variations in speech spectra, leading to improved SSL performance.
- The developed algorithm offers a robust and efficient solution for real-time SSL, suitable for portable, battery-operated devices.
- This work has significant implications for enhancing audio processing in smart loudspeakers and hearing improvement devices.
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Sinusoidal Sources
In homes, the power supplies use sinusoidal sources to provide electricity. These sources generate a voltage that varies sinusoidally...
Sampling Continuous Time Signal
In the...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

