Related Experiment Video
Updated: Sep 18, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Multichannel speech enhancement for automatic speech recognition: a literature review
Zubair Zaland1, Mumtaz Begum Mustafa1, Miss Laiha Mat Kiah2
1Department of Software Engineering, Faculty of Computer Science and Information Technology, Universiti Malaya, Kuala Lumpur, Malaysia.
This review systematically analyzes multichannel speech enhancement (MCSE) for automatic speech recognition (ASR). It identifies effective MCSE methods and highlights challenges in noise generalization for improved ASR performance.
Area of Science:
- Signal Processing
- Artificial Intelligence
- Acoustics
Background:
- Multichannel speech enhancement (MCSE) is vital for robust automatic speech recognition (ASR).
- Existing reviews lack a systematic analysis of MCSE specifically for ASR applications.
- Rapid advancements in MCSE methods, models, and datasets necessitate a comprehensive overview.
Purpose of the Study:
- To systematically review MCSE approaches for ASR systems.
- To analyze the performance of MCSE and ASR across various techniques, models, and noise conditions.
- To discuss current challenges, limitations, and future research directions in MCSE for ASR.
Main Methods:
- Systematic literature review using keyword searches across major electronic databases (Google Scholar, IEEE Xplore, etc.).
- Inclusion and exclusion criteria applied to 240 initial articles, resulting in 40 final experimental articles (23 journals, 17 conferences).
- Backward snowballing and quality assessment were used to finalize the article selection.
Main Results:
- An increasing trend in MCSE for ASR research was observed, with Word Error Rate (WER), Perceptual Evaluation of Speech Quality (PESQ), and Short-Time Objective Intelligence (STOI) as common performance metrics.
- A significant challenge identified is the lack of generality and comparability across MCSE studies, hindering unified solutions for speech recognition noise.
- Key MCSE methods that enhance ASR performance across diverse models, techniques, noise types, and environments were identified.
Conclusions:
- This review provides a comprehensive examination of MCSE and ASR techniques.
- Identified MCSE methods offer potential for improving ASR performance under various conditions.
- Future research should address the generality and comparability issues to develop more robust and unified MCSE solutions for ASR.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Air-entraining Agents

