Related Experiment Video
Updated: Sep 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging multimodal large language model for multimodal sequential recommendation
Zhaoliang Wang1,2, Baisong Liu3, Weiming Huang1
1Faculty of Information Science and Engineering, Ningbo University, Ningbo, 315211, People's Republic of China.
Multimodal large language models (MLLMs) enhance sequential recommendation systems by fusing multimodal features and modeling dynamic user preferences. MLLM-SRec improves recommendation precision and robustness by leveraging MLLMs for better cross-modal understanding.
Area of Science:
- Artificial Intelligence
- Computer Science
- Machine Learning
Background:
- Conventional multimodal recommendation systems struggle with insufficient information exploitation, limited multimodal feature recognition, and ineffective dynamic preference modeling.
- Existing approaches often rely on unimodal data, failing to capture cross-modal preferences and the evolution of user interests in sequential interactions.
- Multimodal large language models (MLLMs) offer advanced cross-modal comprehension and world knowledge, presenting a promising avenue for recommendation system enhancement.
Purpose of the Study:
- To introduce MLLM-SRec, a novel sequential recommendation architecture leveraging MLLMs to address limitations in current multimodal recommendation systems.
- To develop a multimodal feature fusion mechanism for unified item representations, aligning vision and text while mitigating cross-modal differences and noise.
- To design a temporal-aware module for dynamic user preference modeling and integrate it with Chain-of-Thought prompting for effective knowledge transfer.
Main Methods:
- Developed a multimodal feature fusion mechanism using MLLMs to create unified semantic item representations, ensuring semantic alignment between visual and textual data.
- Implemented a temporal-aware user behavior comprehension module to capture the dynamic evolution of user preferences within sequential interaction data.
- Employed supervised fine-tuning combined with multistep Chain-of-Thought prompting to optimize knowledge transfer from pre-trained MLLMs to the recommendation task.
Main Results:
- The proposed MLLM-SRec architecture achieved significant improvements over state-of-the-art baselines across four benchmark datasets.
- The method substantially enhanced the precision of recommendation results.
- MLLM-SRec demonstrated superior robustness and adaptability in multimodal sequential recommendation scenarios.
Conclusions:
- MLLM-SRec effectively addresses the challenges of multimodal feature recognition and dynamic preference modeling in sequential recommendation.
- The architecture validates the significant potential of MLLMs for advancing sequential recommendation tasks by improving multimodal interaction data utilization.
- The findings offer new methodological insights for multimodal sequence recommendation research and highlight the benefits of MLLM integration.
More Related Videos
Related Concept Videos
Language and Cognition
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Associative Learning
Classical conditioning, also known...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...

