Related Experiment Video
Updated: Nov 28, 2025

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention
Published on: December 20, 2024
DMMAN: A two-stage audio-visual fusion framework for sound separation and event localization
Ruihan Hu1, Songbing Zhou1, Zhi Ri Tang2
1Institute of Intelligent Manufacturing, Guangdong Academy of Sciences, Guangdong Key Laboratory of Modern Control Technology, Guangzhou, China.
This study introduces the Deep Multi-Modal Attention Network (DMMAN) to separate and localize sounds in videos. DMMAN effectively handles mixed audio, improving sound source separation and event localization accuracy.
Area of Science:
- Computer Science
- Signal Processing
- Artificial Intelligence
Background:
- Videos contain mixed sounds, making it difficult to distinguish individual sound sources.
- Current methods struggle with separating and localizing sounds in unconstrained video datasets.
Purpose of the Study:
- To develop a model for sound source separation and event localization in videos.
- To address the challenge of mixed audio in multimedia content.
Main Methods:
- Introduced the Deep Multi-Modal Attention Network (DMMAN) for unconstrained video datasets.
- Utilized a multi-modal separator and matching classifier with two-stage fusion of audio-visual features.
- Employed regression and classification losses for model training.
Main Results:
- DMMAN achieved high-quality sound source separation, validated by Signal-to-Distortion Ratio and Signal-to-Interference Ratio metrics.
- The model demonstrated effectiveness in mixed sound scenes previously unaddressed.
- Achieved superior classification accuracy for event localization compared to baseline methods.
Conclusions:
- DMMAN successfully separates sound sources and localizes events in videos.
- The model offers a robust solution for complex audio-visual analysis tasks.
- This approach advances the capabilities of multimedia content understanding.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Auditory Perception
Sound Waves: Interference
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

