Related Experiment Video
Updated: Jul 8, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Zero-Shot Generative Pre-Trained Transformer-Based Multimodal Large Language Models for Pressure Injury
Toshiaki Takahashi1,2, Kengo Miyo2, Nao Tamai1
1Department of Adult Nursing, Graduate School of Medicine, Yokohama City University, Yokohama, Kanagawa, Japan.
Advances in Wound Care
|July 3, 2026
Summary
Generative pre-trained transformer (GPT)-based multimodal large language models (MLLMs) show promise for pressure injury (PI) staging. However, autonomous exact staging is not yet clinically ready, requiring clinician supervision for triage.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Imaging Analysis
- Clinical Decision Support Systems
Background:
- Pressure injuries (PIs) are a significant healthcare concern, necessitating accurate and timely staging for effective management.
- Current methods for PI staging can be subjective and time-consuming, highlighting the need for objective, automated solutions.
- Generative pre-trained transformer (GPT)-based multimodal large language models (MLLMs) offer potential for image analysis and clinical data interpretation.
Purpose of the Study:
- To benchmark zero-shot GPT-based MLLMs for pressure injury staging from photographs.
- To evaluate the impact of prompt strategies, structured outputs, and label granularity on MLLM performance for PI staging.
- To quantify the accuracy and reliability of MLLMs in differentiating PI stages.
Main Methods:
- A retrospective observational benchmark was conducted using 1,091 de-identified PI photographs labeled Stage I-IV.
- Ten model/prompt conditions were evaluated via a standardized API pipeline.
- Performance was assessed using metrics including accuracy, F1 scores, weighted kappa, sensitivity, and specificity for various staging granularities (exact four-class, three-class, skin-break screening, advanced-intervention threshold).
Main Results:
- The best performance for skin-break screening achieved 93.77% accuracy, and for the advanced-intervention threshold, 90.28% accuracy.
- Exact four-class staging yielded the lowest accuracy at 65.08%, indicating challenges with precise differentiation.
- Stage IV undercalling was a safety concern, with one model exhibiting a Stage IV recall of only 8.79% and a false-negative rate of 91.21%.
Conclusions:
- GPT-based MLLMs show potential to assist clinicians in PI triage and prioritization when supervised.
- Autonomous, image-only PI staging using current MLLMs is not yet clinically viable due to accuracy limitations, particularly for advanced stages.
- Further research and development are needed to improve the reliability and safety of MLLMs for clinical PI staging.