Related Experiment Video
Updated: Oct 31, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.7K
Deep Learning Based Real-time Speech Enhancement for Dual-microphone Mobile Phones
Ke Tan1, Xueliang Zhang2, DeLiang Wang3
1Department of Computer Science and Engineering, The Ohio State University, Columbus, OH, 43210-1277 USA.
Summary
This study introduces a new deep learning method for real-time speech enhancement in dual-microphone mobile phones. The system effectively suppresses background noise, improving mobile communication quality.
Area of Science:
- Signal Processing
- Artificial Intelligence
- Mobile Communications
Background:
- Mobile speech communication is often degraded by background noise.
- Speech enhancement systems using microphones are common in mobile phones.
- Existing methods may not be optimal for dual-microphone setups.
Purpose of the Study:
- To propose a novel deep learning approach for real-time speech enhancement.
- To develop an efficient system for dual-microphone mobile phones.
- To improve the quality of speech signals in noisy mobile environments.
Main Methods:
- A densely-connected convolutional recurrent network for dual-channel complex spectral mapping.
- Structured pruning technique for model compression.
- Real-time processing for low-latency and memory efficiency.
Main Results:
- The proposed deep learning approach significantly enhances speech quality.
- The system demonstrates superior performance compared to previous dual-channel methods.
- Outperforms a deep learning-based beamformer in speech enhancement.
Conclusions:
- The novel deep learning approach offers effective real-time speech enhancement for dual-microphone mobile phones.
- Model compression techniques enable efficient and low-latency processing.
- This method represents a significant advancement in mobile communication audio quality.
Related Concept Videos
Design Example
420
The innovation of touch-tone telephony revolutionized the telecommunications industry by replacing the traditional rotary dial with a dual-tone multi-frequency (DTMF) signaling system. This system uses a matrix-style keypad with buttons arranged in four rows and three columns, creating 12 distinct signals each assigned to a pair of frequencies. Each button press results in a simultaneous generation of two sinusoidal tones – one from a low-frequency group (697 to 941 Hz) and one from a...
420
Perceiving Loudness, Pitch, and Location
566
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
566
Air-entraining Agents
144
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
144

