Related Experiment Videos
Exposure-aware multimodal engagement prediction for short-video platforms: an explainable human-media interaction
1The Education University of Hong Kong, Hong Kong, Hong Kong SAR, China.
Abstract:
Short-video recommendation feeds are increasingly important human-media interaction systems. However, the observable pathways through which algorithmic exposure is associated with multimodal content and measurable user engagement remain insufficiently explained. Existing recommender-system studies usually optimize ranking accuracy or watch-time objectives. Communication-oriented studies often examine engagement without modeling the exposure conditions under which interaction data are generated. This paper develops an explainable exposure-aware multimodal framework for short-video engagement prediction and observable algorithmic amplification measurement. The framework combines the Algorithmic Amplification Gap (AAG), random-exposure baselines, multi-target engagement modeling, exposure-gated multimodal fusion, and interpretable feature attribution. KuaiRand-Pure is used as the main exposure-aware dataset, containing 27,285 users, 7,551 short videos, and 1,436,609 user-video interaction records. KuaiRec is used as an auxiliary robustness benchmark. MicroLens is introduced as an external raw-multimodal validation dataset with titles, cover images, audio, and full-length micro-videos. Across experiments, exposure proxies and temporally separated early-window feedback signals provide substantial explanatory power. Textual features, semantic/category proxy features, temporal/retention-behavioral proxy features, and raw multimodal representations remain important for target-specific engagement patterns. The proposed Exposure-Gated Multimodal Fusion component achieves statistically reliable but incremental gains over concatenation-based fusion. The main contribution is therefore an auditable human-media interaction protocol for examining how observable algorithmic visibility, multimodal content, and user feedback jointly shape short-video engagement.
Related Concept Videos
Introducing Social Perception
Facial Feedback Hypothesis