Related Experiment Video
Updated: Aug 7, 2026

04:32
Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention
Published on: December 20, 2024
Dual-path magnitude-phase learning with bidirectional cross-attention for bone-conducted speech restoration
Dianzhe Ding1,2, Sichen Liu3, Junfeng Li4
1Laboratory of Noise and Audio Research, Institute of Acoustics, Chinese Academy of Sciences, Beijing 100190, China.
The Journal of the Acoustical Society of America
|August 5, 2026
Summary
This study introduces a novel dual-path U-Net model for restoring air-conducted speech from bone-conducted speech. The model effectively enhances speech quality in low signal-to-noise ratios using magnitude and phase information.
Area of Science:
- Signal Processing
- Artificial Intelligence
- Speech Technology
Background:
- Bone-conducted speech offers noise immunity but limited bandwidth.
- Restoring air-conducted speech from bone-conducted speech is crucial for low signal-to-noise ratio environments.
- Existing methods for speech restoration have limitations.
Purpose of the Study:
- To develop an advanced model for converting bone-conducted speech to high-quality air-conducted speech.
- To leverage both magnitude and phase information from bone-conducted speech spectrograms.
- To improve speech restoration performance compared to existing techniques.
Main Methods:
- A dual-path frequency-domain U-Net (UNet) model was designed.
- Dual-path convolution modules extract magnitude and phase features separately.
- A bidirectional cross-attention module integrates information from both features.
- A frequency-weighted spectrogram loss function was employed for training.
Main Results:
- The proposed model significantly outperforms existing bone-conducted speech restoration methods.
- The model achieves high-quality air-conducted speech restoration.
- The model demonstrates efficiency with fewer parameters and operations compared to competitors.
Conclusions:
- The dual-path UNet model effectively restores air-conducted speech from bone-conducted speech.
- Jointly utilizing magnitude and phase information enhances restoration quality.
- The proposed method offers a computationally efficient and high-performing solution for speech enhancement.