Related Experiment Video
Updated: Jun 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
AI-based burn image assessment: Reliability and clinical error patterns of multimodal large language models in a
Ibrahim Güler1, Armin Kraus2, Gerrit Grieb3
1Department of Plastic, Aesthetic and Hand Surgery, University Hospital Magdeburg, Leipziger Strasse 44, 39120 Magdeburg, Germany; Department of Health Management, Friedrich-Alexander-Universität Erlangen-Nürnberg FAU, Lange Gasse 20, 90403 Nürnberg, Germany.
Multimodal large language models (MLLMs) show inconsistent performance in assessing burn depth and total body surface area (TBSA), with accuracy varying significantly and reliability issues hindering clinical application.
Area of Science:
- Medical Imaging Analysis
- Artificial Intelligence in Healthcare
- Burn Management
Background:
- Accurate burn depth and Total Body Surface Area (TBSA) assessment is crucial for effective clinical management.
- Current assessment methods are subjective and suffer from interobserver variability.
- The potential of Multimodal Large Language Models (MLLMs) in medical image analysis, specifically for burns, is largely unexplored.
Purpose of the Study:
- To evaluate the reliability and accuracy of four leading MLLMs in assessing burn depth and TBSA from clinical photographs.
- To identify performance variations and biases among different MLLM systems.
- To determine if MLLMs can provide consistent and dependable burn assessments for clinical use.
Main Methods:
- Fifty clinical burn photographs were analyzed by four MLLMs (GPT-5.4 Pro, Grok 4.1, Gemini 3.1 Pro, Claude Opus 4.6).
- A repeated-inference design with five independent runs per model was employed.
- Burn depth (numeric/text) and TBSA (ordinal) were assessed, with performance measured by accuracy and inter-run reliability (Fleiss' κ).
Main Results:
- MLLM accuracy for burn depth ranged from 34.0% to 76.4%, and for TBSA from 32.8% to 68.4%.
- Inter-run reliability varied widely (κ = 0.171 to 0.916), with no model achieving both high accuracy and reliability.
- Models exhibited biases, including overestimation of burn depth and inconsistent error patterns, despite high internal consistency between output formats.
Conclusions:
- Current MLLMs demonstrate significant variability and inconsistency in burn assessment, rendering them unreliable for clinical decision-making.
- Stochastic response instability in MLLMs, not apparent in single-query evaluations, poses a fundamental limitation.
- Further development is required to address reliability and bias issues before MLLMs can be integrated into burn care workflows.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy