Related Experiment Videos
AdaCompVL: Adaptive Compression of Spatiotemporal and Cross-Modal Redundancy for Efficient Video-Language Learning
Abstract:
Multimodal large language models (MLLMs) have recently extended from static image understanding to video comprehension, but representing videos as frame-level token sequences incurs substantial computational overhead. Existing visual token compression methods typically rely on uniform sampling or single-dimension redundancy modeling, making them less effective across diverse scene dynamics and textual inputs. To address this issue, we propose AdaCompVL (Adaptive Compression for Video-Language), a training-free adaptive compression framework that considers both scene dynamics and semantic relevance to efficiently reduce redundant visual tokens. Specifically, we introduce a dynamic-aware switching mechanism that leverages adjacent-frame similarity to distinguish between low-and high-dynamic scenarios. For low-dynamic scenes, temporal redundancy pruning removes minimally changing tokens across frames, whereas for high-dynamic scenes, a saliency-guided spatial pruning strategy retains visually informative intra-frame tokens. To further reduce cross-modal redundancy, a mutual information-based module preserves only the visual tokens most relevant to the textual input, yielding a compact yet semantically meaningful video-language representation. Extensive experiments on multiple video-language benchmarks show that AdaCompVL removes 85% of visual tokens while retaining 99% of the original performance, achieving up to a 1.75× inference speedup. These results demonstrate that AdaCompVL produces compact and semantically relevant video representations for efficient video-language modeling. Our code is publicly available at https: //github.com/JianXin-M/AdaCompVL.
Related Concept Videos
Associative Learning
Classical conditioning, also known...
Learning Disabilities
Dyslexia
Dyslexia is a...
Chunking and Rehearsal in Sensory Memory
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Language and Cognition