Related Experiment Video
Updated: Sep 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Interleaved Prompt Generation for Continual Instruction Tuning of Multimodal Large Language Model
Abstract:
Multimodal Large Language Models (MLLMs), pre-trained on vast image-text datasets, excel in zero-shot reasoning but face catastrophic forgetting when adapting to sequential tasks in Continual Instruction Tuning (CIT). While lightweight prompt-based methods offer a parameter-efficient strategy for sequential adaptation, these approaches face a key limitation: conventional global prefix prompts, owing to the causal masking inherent in decoder-based MLLMs, cannot directly interact with visual tokens. This constraint limits the interaction between task instructions and visual semantics, restricting the ability of visual information to serve as stable, discriminative anchors for knowledge retention. Moreover, static prompt representations struggle to capture intra-task instance variations, limiting fine-grained adaptation to diverse multimodal inputs. To address these challenges, we propose Interleaved Prompt Generation (IPG), a novel framework that jointly addresses prompt placement, instance-conditioned generation, and prompt knowledge organization for continual multimodal adaptation, with three complementary designs: (1) an interleaved prompt structure that inserts learnable tokens within the visual token sequence, enabling direct prompt-visual interaction despite decoder causal constraints; (2) instance-level prompt generation conditioned on frozen CLIP features, which dynamically generates instance-aware prompts by adaptively modulating shared representations through Feature-wise Linear Modulation; and (3) a stochastic mixture of Prompt Generators that maintains diverse prompt experts and employs Dirichlet-sampled routing to selectively reuse and fuse historical prompt knowledge across sequential tasks. Our method achieves state-of-the-art performance in continual instruction tuning on split-VQAv2 and the CoIN benchmark, validating its efficacy in adapting MLLMs to sequential multimodal tasks.
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Elaborative Rehearsals
The effectiveness of...
Components of Language
Language and Cognition
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Purposive Learning