Related Experiment Video
Updated: Sep 27, 2026

Rapid Setup of Tissue Microarray and Tiled Area Imaging on the Multiplexed Ion Beam Imaging Microscope Using the Tile/SED/Array Interface
Published on: September 15, 2023
AI-Based Spatial Burn Assessment with MLLMs: Body Region, 3 × 3 Grid Classification and Burn Instance Counting from
Ibrahim Güler1,2, Armin Kraus1, Gerrit Grieb3,4
1Department of Plastic, Aesthetic and Hand Surgery, Otto-von-Guericke University, 39120 Magdeburg, Germany.
Abstract:
Background: Burn photography is a stringent testbed for visual artificial intelligence (AI) in medical imaging, since it permits quantitative and qualitative clinically relevant parameters to be elicited. It is therefore a useful proxy for current vision AI. Reducing a photograph to a semantic segmentation removes color, texture and anatomy while providing pixel-exact boundaries. How spatial assessment responds, and how far instance-level structure can be recovered from a map that does not encode it, is unknown. Methods: Three state-of-the-art multimodal large language models (MLLMs) each assessed 153 burn photographs five times under three input conditions: the photograph alone (IMAGE), its segmentation mask alone (MASK), or both combined (BOTH). Four items were elicited per image: burn presence, the burned body region, per-cell classification of a 3 × 3 grid, and the number of separate burns, a surrogate for instance-level recognition. Each model was compared across conditions, paired on the same images. Results: Body-region accuracy was 97.1-100.0% under IMAGE, 47.5-54.0% under MASK and 50.2-99.9% under BOTH. Grid-cell accuracy was 77.7-82.3%, 77.9-96.3% and 90.9-98.4%. On the burn count, exact-match was 60.5-66.3%, 61.3-96.6% and 84.7-97.4%, and no model significantly beat the trivial classifier under IMAGE, two of three did under MASK and all three under BOTH. Across 18 paired comparisons of BOTH against a single channel, BOTH was better in 11, worse in three and indistinguishable in four. Conclusions: Combining the two inputs improved spatial assessment in most comparisons and degraded it in others, consistently across neither models nor tasks. Additional visual information therefore does not deterministically improve model performance, and disruption remains possible.