Related Experiment Video
Updated: Sep 17, 2025

Enhancing Electrode Location Assessment in Cochlear Implantation via Computed Tomography Image Fusion
Published on: January 17, 2025
End-to-end feature fusion for jointly optimized speech enhancement and automatic speech recognition
Mohamed Medani1, Nasir Saleem2, Fethi Fkih3
1Applied College of Muhayel Aseer, King Khalid University, Abha, 62529, Saudi Arabia.
This study introduces a novel speech enhancement (SE) model that dynamically fuses enhanced and noisy speech features. This approach significantly improves speech quality and reduces errors in real-time automatic speech recognition (ASR) systems.
Area of Science:
- Signal Processing
- Artificial Intelligence
- Machine Learning
Background:
- Real-time speech enhancement (SE) and automatic speech recognition (ASR) are critical for clear communication and accurate transcription.
- Traditional SE methods can distort speech, negatively impacting downstream ASR performance.
- Existing joint SE-ASR models often use only enhanced features, potentially losing valuable information.
Purpose of the Study:
- To develop an SE network that suppresses noise while minimizing speech distortion.
- To propose a dynamic fusion approach integrating enhanced and raw noisy speech features for improved ASR.
- To create a joint training framework for robust end-to-end ASR.
Main Methods:
- An attentional codec model with a causal attention mechanism for SE.
- A modified Gated Recurrent Unit (GRU) in the SE network, using an attention-based rectified linear unit (AReLU).
- A GRU-based fusion network combining enhanced and raw noisy features, feeding into an ASR system.
Main Results:
- The proposed SE model achieved superior speech quality, intelligibility, and noise suppression compared to baselines in matched and unmatched conditions.
- Significant improvements in STOI (19.81%) and PESQ (28.97%) were observed in matched conditions.
- The joint training framework reduced character error rate (CER) from 32.99% to 13.52% for ASR.
Conclusions:
- The dynamic fusion approach effectively mitigates speech distortions and preserves crucial details from noisy signals.
- The proposed SE and joint training framework significantly enhance the robustness and accuracy of real-time ASR systems.
- This integrated approach offers a promising solution for challenging acoustic environments.
More Related Videos
Related Concept Videos
Amplifying Signals via Enzymatic Cascade
Improving Translational Accuracy
Extraction: Advanced Methods
Double Resonance Techniques: Overview
Spin decoupling is usually achieved by...
Pre-mRNA Processing: Modification of pre-mRNA Ends
Once about 20-40 ribonucleotides have been joined together by RNA polymerase, a group of enzymes adds a cap to the 5' end of the growing transcript. In this process, a 5' phosphate is replaced by modified guanosine that has a methyl group attached (7-methyl guanosine). This 5' cap helps...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....

