Related Experiment Video
Updated: Jun 19, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Leveraging VLMs for MUDA: Category-specific prompt with multi-modal interactive LoRA.
Jianing Yang1, Xihuai He1, Xueqiong Li1
1College of Computer Science and Technology, National University of Defense Technology, Changsha, 410000, China.
This study introduces a novel CLIP-model framework for Multi-Source Unsupervised Domain Adaptation (MUDA). It effectively adapts models to new domains by integrating category-specific prompts and multimodal Low-Rank adaptation, significantly improving performance.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Multi-Source Unsupervised Domain Adaptation (MUDA) aims to train models using labeled data from multiple sources and unlabeled target data.
- Existing MUDA methods lack effective integration with emerging pre-trained Visual Language Models (VLMs).
Purpose of the Study:
- To develop a novel framework for MUDA leveraging CLIP-based Visual Language Models.
- To address the limitations of current methods in adapting models to target domains using VLMs.
Main Methods:
- A CLIP-model-based framework integrating category-specific prompts and multimodal Low-Rank (LoRA) matrix adaptation.
- Utilizing learnable, class-specific prompts for shared knowledge extraction.
- Employing multimodal LoRA for domain-specific knowledge acquisition and a modality interaction mechanism.
Main Results:
- The proposed method demonstrates significant improvements on standard image classification benchmark datasets.
- Successful adaptation to target domains using a combination of shared and domain-specific knowledge.
Conclusions:
- The novel CLIP-model-based framework offers an effective solution for MUDA.
- The integration of category-specific prompts and multimodal LoRA advances VLM-based domain adaptation techniques.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Response Surface Methodology
The process of RSM involves several key steps:
Sensory Modalities
General senses refer to the broad category of sensory information detected by receptors in the body and can be further grouped into somatic and visceral senses. Somatic sensations include touch, pressure, temperature, and pain and are essential for navigating our environment and...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...
Lagrange Multipliers: Problem Solving
Methods of Medium Optimization