Related Experiment Video
Updated: May 26, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
Yogesh Kulkarni1, Pooyan Fazli1
1Arizona State University.
VideoPASTA (Preference Alignment with Spatio-Temporal-Cross Frame Adversaries) improves video-language models by training them to identify flawed video representations. This approach enhances understanding of spatial and temporal details without human annotation.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Video-language models (Video-LLMs) demonstrate proficiency in video comprehension but exhibit weaknesses in spatial reasoning, temporal sequencing, and cross-frame consistency.
- Existing models often require extensive pretraining or architectural changes to improve performance.
Purpose of the Study:
- To introduce VideoPASTA (Preference Alignment with Spatio-Temporal-Cross Frame Adversaries), a novel framework designed to enhance Video-LLMs.
- To improve the ability of Video-LLMs to understand complex spatial and temporal dynamics within videos.
Main Methods:
- VideoPASTA employs targeted preference optimization, training Video-LLMs to differentiate correct video representations from adversarial examples that violate spatial, temporal, or cross-frame relationships.
- The framework utilizes Direct Preference Optimization with a limited dataset of 7,020 preference pairs and 32-frame sampling.
Main Results:
- VideoPASTA significantly enhances the performance of various state-of-the-art Video-LLMs across multiple benchmarks, including LongVideoBench (+3.8%), VideoMME (+4.1%), and MVBench (+4.0%).
- The approach is model-agnostic and achieves substantial improvements without requiring human annotation or captioning.
- Models trained with VideoPASTA demonstrate improved capture of fine-grained spatial details and long-range temporal dynamics.
Conclusions:
- Targeted preference alignment, as implemented in VideoPASTA, is an effective strategy for addressing core challenges in video-language understanding.
- VideoPASTA offers a scalable, plug-and-play solution that seamlessly integrates with existing Video-LLMs, enhancing their capabilities without altering their fundamental architecture.
Related Concept Videos
Impression Management Techniques IV: Altercasting
Kendall's Coefficient of Concordance
The Anchoring-and-Adjustment Heuristic
Impression Management Techniques III: Aligning Actions
Wilcoxon Signed-Ranks Test for Matched Pairs
Friedman Two-way Analysis of Variance by Ranks