Related Experiment Video
Updated: May 10, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Parameter-Efficient Adaptation of Large Vision-Language Models for Video Memorability Prediction
Iván Martín-Fernández1, Sergio Esteban-Romero1, Fernando Fernández-Martínez1
1Grupo de Tecnología del Habla y Aprendizaje Automático (THAU Group), Information Processing and Telecommunications Center, E.T.S.I. de Telecomunicación, Universidad Politécnica de Madrid (UPM), 28040 Madrid, Spain.
This study enhances video memorability prediction by adapting Large Vision-Language Models (LVLMs) using Quantized Low-Rank Adaptation (QLoRA). The fine-tuned Qwen-VL model achieved state-of-the-art results, improving media analysis and generation.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Multimedia Analysis
Background:
- Accurate video memorability modeling is crucial for efficient media retrieval, classification, and generation.
- Strong correlation exists between video visual semantics and memorability, necessitating advanced visual comprehension.
- Large Vision-Language Models (LVLMs) excel at high-level semantic understanding due to extensive multimodal pre-training.
Purpose of the Study:
- To leverage LVLMs for video memorability prediction.
- To explore efficient adaptation techniques for LVLMs in memorability modeling.
- To investigate the impact of LoRA hyperparameters on memorability prediction performance.
Main Methods:
- Fine-tuning the Qwen-VL model using the Quantized Low-Rank Adaptation (QLoRA) technique.
- Utilizing memorability-related data from the Memento10k dataset for adaptation.
- Transforming Qwen-VL into a memorability score regressor.
- Optimizing LoRA hyperparameters (rank and alpha) via 5-Fold Cross-Validation.
Main Results:
- Achieved a state-of-the-art Spearman Rank Correlation Coefficient (SRCC) of 0.744 on the Memento10k dataset.
- Demonstrated the effectiveness of QLoRA for adapting LVLMs to memorability prediction.
- Identified optimal LoRA hyperparameters for improved performance.
Conclusions:
- This work significantly advances video memorability modeling through LVLMs and efficient adaptation.
- The proposed methodology offers a robust approach for predicting video memorability.
- High-level semantic understanding is key to accurate video memorability prediction.
More Related Videos
Related Concept Videos
Associative Learning
Classical conditioning, also known...
Improving Translational Accuracy
Chunking and Rehearsal in Sensory Memory
Vision
Elaborative Rehearsals
The effectiveness of...

