Related Experiment Video
Updated: Jun 23, 2026

10:42
A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
Artificial intelligence-based decision support in simulated free flap Re-exploration for head and neck
Sebastian Holm1,2, Juan E Berner3, Ann-Charlott Docherty Skogh1,4
1Department of Reconstructive Plastic Surgery, Karolinska University hospital, Stockholm, Sweden.
JPRAS Open
|June 22, 2026
Summary
Large language models show promise in recognizing free flap compromise in head and neck reconstruction. However, differences in management recommendations and clarity necessitate cautious use in microsurgery.
Area of Science:
- Microsurgery
- Artificial Intelligence
- Head and Neck Reconstruction
Background:
- Free flap compromise is a critical complication in head and neck reconstruction.
- Prompt recognition and management are essential for successful outcomes.
- The role of large language models (LLMs) in supporting microsurgical decision-making for flap compromise is not well-defined.
Purpose of the Study:
- To evaluate the performance of three LLMs (ChatGPT, Gemini, Copilot) in simulated head and neck free flap compromise scenarios.
- To compare LLM accuracy, management recommendations, and potential for harm against expert microsurgeon ratings.
- To assess LLM utility in both intraoperative and postoperative flap compromise situations.
Main Methods:
- A simulated case study using 30 standardized head and neck free flap compromise scenarios (15 intraoperative, 15 postoperative).
- Three LLMs (ChatGPT, Gemini, Copilot) were prompted for diagnosis, mechanism, and management.
- Four consultant microsurgeons rated LLM responses on clinical accuracy, management, explanation, confidence, and potential for harm.
Main Results:
- All LLMs demonstrated high clinical accuracy in recognizing flap compromise.
- Gemini outperformed ChatGPT and Copilot in management recommendations and explanation clarity for postoperative cases.
- Potentially harmful recommendations were infrequent but varied between models, particularly in postoperative scenarios.
Conclusions:
- LLMs show significant alignment with expert recognition of flap compromise.
- Differences in management advice and explanatory clarity, especially in postoperative settings, warrant careful consideration.
- Current LLMs should be used cautiously and do not currently support independent clinical use in microsurgical care.