Related Experiment Video
Updated: Aug 15, 2026

05:48
Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Robust Audio-Visual Question Answering with Missing Modality in Training and Testing
IEEE Transactions on Pattern Analysis and Machine Intelligence
|August 13, 2026
Summary
This study introduces Training-time Modality-Missing Audio-Visual Question Answering (TM-AVQA) to handle missing audio or visual data. The proposed framework effectively reconstructs missing modalities, improving AVQA performance even with incomplete data.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Machine Learning
Background:
- Audio-Visual Question Answering (AVQA) typically requires complete audio and visual data for reasoning over dynamic scenes.
- Real-world scenarios often involve missing modalities due to signal loss or privacy concerns, challenging existing AVQA methods.
- A new setting, Training-time Modality-Missing AVQA (TM-AVQA), is proposed to address AVQA with potentially incomplete data during training and testing.
Purpose of the Study:
- To develop a robust AVQA framework capable of handling missing audio or visual modalities during both training and inference.
- To introduce novel techniques for cross-modal reconstruction and dependency modeling under modality-missing constraints.
- To evaluate the effectiveness of the proposed methods on established AVQA datasets adapted for modality-missing scenarios.
Main Methods:
- A two-stage AVQA framework is proposed, incorporating a reconstruction network and an answer prediction stage.
- Stage-I features a reconstruction network with Dense Temporal-scale Reconstruction (DTR) and Multimodal Dependency Modeling (MDM) modules.
- Novel learning objectives, Cross-Modal Relation-based Contrastive Learning (CMR-CL) and Cross-Sample Relation-based Pseudo-label Learning (CSR-PL), are adapted to supervise reconstruction with missing modalities.
Main Results:
- The proposed TM-AVQA framework demonstrated consistent and robust improvements across various AVQA backbones and missing-modality conditions.
- Experiments on modified MUSIC-AVQA, MUSIC-AVQA-R, and AVQA datasets confirmed the effectiveness of the approach.
- Ablation studies validated the contributions of task-specific designs, reconstruction quality, and efficiency.
Conclusions:
- The developed framework effectively addresses the challenge of missing modalities in AVQA, outperforming existing baselines.
- The proposed reconstruction and learning strategies enable robust performance even when one modality is entirely absent.
- The approach shows potential for transfer learning to related audio-visual recognition tasks.
Related Concept Videos
Sensory Modalities
Sensation typically is the process by which the sensory receptors and sense organs detect stimuli from the internal and external environment and transmit this information to the central nervous system for processing.
General senses refer to the broad category of sensory information detected by receptors in the body and can be further grouped into somatic and visceral senses. Somatic sensations include touch, pressure, temperature, and pain and are essential for navigating our environment and...
General senses refer to the broad category of sensory information detected by receptors in the body and can be further grouped into somatic and visceral senses. Somatic sensations include touch, pressure, temperature, and pain and are essential for navigating our environment and...
Auditory Pathway
Auditory pathways constitute the complex neural circuits responsible for transmitting and interpreting auditory information from the peripheral auditory system to the brain. Sound waves are initially captured by the outer ear, funneled through the ear canal, and reach the tympanic membrane (eardrum). These vibrations are transmitted via the middle ear's ossicles to the inner ear's cochlea.
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking the...
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking the...
Auditory Perception
The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the cochlea, a...
Non-Verbal Cues
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
Chunking and Rehearsal in Sensory Memory
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of information more...
