Related Experiment Video
Updated: Jul 4, 2025

Investigating the Three-dimensional Flow Separation Induced by a Model Vocal Fold Polyp
Published on: February 3, 2014
3 directional Inception-ResUNet: Deep spatial feature learning for multichannel singing voice separation with
DaDong Wang1, Jie Wang1, MingChen Sun2
1School of Mathematics and Computer Science, Jilin Normal University, Siping, Jilin, China.
This study introduces a novel 3D Inception-ResUNet model for robot singing voice separation, significantly improving performance by utilizing spatial and spectral information. The multi-objective training approach achieved an average 11.55 dB NSDR, outperforming existing methods.
Area of Science:
- Robotics
- Signal Processing
- Artificial Intelligence
Background:
- Humanoid robots face challenges in interpreting complex auditory signals, such as mixed singing voices, music, and noise.
- Acoustic signals perceived by robots are often distorted, attenuated, and reverberated, complicating voice separation tasks.
Purpose of the Study:
- To develop an advanced singing voice separation model for humanoid robots.
- To enhance the robot's ability to interpret ambiguous auditory signals in noisy environments.
Main Methods:
- Utilized a 3D Inception-ResUNet architecture within a U-shaped network to process spectrograms.
- Employed multi-objective training with magnitude consistency loss, phase consistency loss, and magnitude correlation consistency loss.
- Synthesized a 10-channel dataset using NAO robots and the MIR-1K dataset for model training.
Main Results:
- The proposed model achieved an average Normalized Source-to-Distortion Ratio (NSDR) of 11.55 dB on the test dataset.
- Demonstrated superior performance compared to existing comparison models in singing voice separation tasks.
- The multi-objective training strategy effectively improved the utilization of spatial and spectral information.
Conclusions:
- The 3D Inception-ResUNet model offers a significant advancement in robot-based singing voice separation.
- Multi-objective training is crucial for enhancing the accuracy and robustness of auditory signal interpretation in robots.
- This research contributes to more sophisticated human-robot interaction through improved audio processing capabilities.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Double Resonance Techniques: Overview
Spin decoupling is usually achieved by...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....

