AI-based burn image assessment: Reliability and clinical error patterns of multimodal large language models in a

Ibrahim Güler1, Armin Kraus2, Gerrit Grieb3

  • 1Department of Plastic, Aesthetic and Hand Surgery, University Hospital Magdeburg, Leipziger Strasse 44, 39120 Magdeburg, Germany; Department of Health Management, Friedrich-Alexander-Universität Erlangen-Nürnberg FAU, Lange Gasse 20, 90403 Nürnberg, Germany.

Summary

Multimodal large language models (MLLMs) show inconsistent performance in assessing burn depth and total body surface area (TBSA), with accuracy varying significantly and reliability issues hindering clinical application.