Related Experiment Video
Updated: Jan 13, 2026

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
AFCLNet: An Attention and Feature-Consistency-Loss-Based Multi-Task Learning Network for Affective Matching
Zhibin Su1,2,3, Jinyu Liu2,3, Luyue Zhang2,3
1State Key Laboratory of Media Convergence and Communication, Beijing 100024, China.
Abstract:
Emotion matching prediction between music and video segments is essential for intelligent mobile sensing systems, where multimodal affective cues collected from smart devices must be jointly analyzed for context-aware media understanding. However, traditional approaches relying on single-modality feature extraction struggle to capture complex cross-modal dependencies, resulting in a gap between low-level audiovisual signals and high-level affective semantics. To address these challenges, a dual-driven framework that integrates perceptual characteristics with objective feature representations is proposed for audiovisual affective matching prediction. The framework incorporates fine-grained affective states of audiovisual data to better characterize cross-modal correlations from an emotional distribution perspective. Moreover, a decoupled Deep Canonical Correlation Analysis approach is developed, incorporating discriminative sample-pairing criteria (matched/mismatched data discrimination) and separate modality-specific component extractors, which dynamically refine the feature projection space. To further enhance multimodal feature interaction, an Attention and Feature-Consistency-Loss-Based Multi-Task Learning Network is proposed. In addition, a feature-consistency loss function is introduced to impose joint constraints across dual semantic embeddings, ensuring both affective consistency and matching accuracy. Experiments on a self-collected benchmark dataset demonstrate that the proposed method achieves a mean absolute error of 0.109 in music-video matching score prediction, significantly outperforming existing approaches.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Associative Learning
Classical conditioning, also known...
Facial Feedback Hypothesis
