Related Experiment Video
Updated: Jan 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Optimizing multimodal models for medical visual question answering: A comparative study of LoRA and AdaLoRA on
Zahra Rezaei1, Sara Safi Samghabadi1, Yaser Mike Banad1
1School of Electrical and Computer Engineering, University of Oklahoma, Norman, OK, USA, 73019.
Abstract:
Medical Visual Question Answering (Med-VQA) empowers AI systems to interpret medical images and respond to clinical queries, enhancing diagnostic precision and decision-making in resource-limited clinical settings. This study evaluates parameter-efficient fine-tuning (PEFT) techniques on leading multimodal models: Idefics3-8B-Llama3, idefics2-8b, LLaVA-1.5-7b, Qwen2-VL-7B-Instruct, and Llama-3.2-11B-Vision-Instruct to optimize their performance for Med-VQA tasks. The key innovation lies in employing Low-Rank Adaptation (LoRA) and Adaptive Low-Rank Adaptation (AdaLoRA), which achieve significant computational efficiency, making these models viable for clinical deployment by reducing resource demands while maintaining high accuracy. Results show that on the VQA-RAD dataset, Idefics3-8B-Llama3 attained 90 % accuracy with AdaLoRA and 86 % with LoRA, while on the SLAKE-VQA dataset, both Idefics3-8B-Llama3 and idefics2-8b reached 93 % accuracy with LoRA, demonstrating their suitability for clinical applications. The study followed a systematic pipeline: selecting and preprocessing the VQA-RAD and SLAKE-VQA datasets, fine-tuning the models using LoRA and AdaLoRA with hyperparameter optimization, and rigorously evaluating their performance. This research establishes new benchmarks and provides actionable insights for integrating AI into medical image analysis, advancing clinical decision support systems.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy