Related Experiment Videos
SDPT: Synchronous Dual Prompt Tuning for Visual-Language Pre-trained Models
Abstract:
Prompt tuning methods use learnable tokens for parameter-efficient downstream adaptation on large pre-trained models. However, for dual-modal visual-language pre-trained models (VLPMs), existing prompt tuning methods overlook the preservation of pre-trained text-image alignment during fine-tuning. To address this issue, we propose Synchronous Dual Prompt Tuning (SDPT). SDPT initializes a single set of learnable unified prototype tokens in the established modal aligning space to represent the aligned semantics of text and image modalities for downstream tasks. Furthermore, SDPT establishes inverse linear projections, whose projection matrices need no training, to embed the information of learnable unified prototype tokens into the input space of different modalities. The inverse linear projections allow the unified prototype token to synchronously represent the two modalities and enable SDPT to share the unified semantics of text and image for downstream tasks across different modal prompts. Experimental results demonstrate that SDPT assists VLPMs to achieve superior outcomes with only 0.04% of model parameters for training across various scenarios, outperforming other single- or dual-modal methods. The code is released at github/SDPT.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Parallel Processing
Multi-input and Multi-variable systems
In the absence of...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...