Related Experiment Video
Updated: Jun 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance, safety, and limitations of multimodal large language models in wound image assessment
Sirin Apichonbancha1, Nutcha Yodrabum1, Sasima Tongsai2
1Division of Plastic Surgery, Department of Surgery, Faculty of Medicine Siriraj Hospital, Mahidol University, Bangkok, Thailand.
Abstract:
Accurate visual assessment of acute, chronic, and surgical wounds is fundamental to clinical decision-making and is increasingly performed using digital wound photographs in routine and remote care. However, interpretation of wound images remains subjective and variable across clinicians, and the rapid emergence of vision-capable large language models (LLMs) has led to their informal use for wound description and clinical interpretation despite limited task-specific validation. To address this gap, we compared three advanced vision-capable LLMs (ChatGPT-4o, Claude 3.5 Sonnet, and Gemini Advanced) for structured wound image assessment using standardized clinical frameworks. From 1,200 clinical wound photographs obtained during routine care, 450 images (150 acute, 150 chronic, and 150 surgical) were randomly selected and independently reviewed by three expert clinicians to establish expert consensus reference standards. Each model received identical prompts incorporating the Bates-Jensen Wound Assessment Tool (BWAT) and the TIMES framework, and outputs were evaluated for diagnostic accuracy, assessment quality, treatment recommendation appropriateness, safety, understandability, and actionability; informational quality was assessed using DISCERN and PEMAT-P. ChatGPT-4o achieved the highest accuracy for diagnosis (51.3%), clinical conclusion (52.4%), and treatment recommendations (53.6%), while Claude performed best for wound sizing (72.0%) and urgency determination (70.9%). Gemini showed substantial limitations, with non-response in 67-68% across several assessment domains and the lowest clinical and safety performance. Overall, ChatGPT-4o most consistently generated accurate, structured, and clinically aligned wound assessments. However, the findings also highlight important limitations in reliability, safety, and real-world applicability, underscoring the need for further validation and explicit safeguards before routine clinical integration.