大视觉语言模型的参数效率适应用于视频记忆性预测
Iván Martín-Fernández1, Sergio Esteban-Romero1, Fernando Fernández-Martínez1
1Grupo de Tecnología del Habla y Aprendizaje Automático (THAU Group), Information Processing and Telecommunications Center, E.T.S.I. de Telecomunicación, Universidad Politécnica de Madrid (UPM), 28040 Madrid, Spain.
Sensors (Basel, Switzerland)
|April 28, 2025
概括
这项研究通过使用量化低等级适应 (QLoRA) 调整大视觉语言模型 (LVLMs) 来增强视频记忆力的预测. 微调的Qwen-VL模型取得了最先进的结果,改善了媒体分析和生成.
科学领域:
- 人工智能的人工智能
- 计算机视觉 计算机视觉
- 多媒体分析分析.
背景情况:
- 准确的视频可记性建模对于高效的媒体检索,分类和生成至关重要.
- 视频视觉语义和记忆能力之间存在强烈的相关性,需要先进的视觉理解.
- 大视觉语言模型 (LVLMs) 由于广泛的多式模式预训练,在高水平的语义理解方面表现出色.
研究的目的:
- 为了利用LVLMs进行视频记忆性预测.
- 在记忆能力建模中探索LVLM的高效适应技术.
- 调查LoRA超参数对记忆力预测性能的影响.
主要方法:
- 使用量子化低等级适应 (QLoRA) 技术微调Qwen-VL模型.
- 使用Memento10k数据集中的与记忆能力相关的数据进行适应.
- 将Qwen-VL转换为一个记忆度得分回归器.
- 通过5倍交叉验证优化LoRA超参数 (等级和alpha).
主要成果:
- 在Memento10k数据集上获得了0.744的最先进的斯皮尔曼等级相关系数 (SRCC).
- 证明了QLoRA在适应LVLM与记忆力预测方面的有效性.
- 确定了最佳的LoRA超参数,以提高性能.
结论:
- 这项工作通过LVLMs和高效适应显著推进了视频记忆能力建模.
- 拟议的方法提供了一个可靠的方法来预测视频记忆力.
- 高层次的语义理解是准确预测视频记忆力的关键.
更多相关视频
相关概念视频
Associative Learning
236
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
236
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Chunking and Rehearsal in Sensory Memory
107
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
107
Vision
52.2K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.2K
Elaborative Rehearsals
58
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
58


