Related Experiment Video
Updated: Feb 7, 2026

Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
Large Language Models and Surgical Decision-Making: Evaluation of Generative Unimodal AI in Facial Traumatology
Simone Benedetti1, Andrea Frosolini1, Lisa Catarzi1
1Maxillofacial Surgery Unit, Department of Medical Biotechnologies, University of Siena, Siena, Italy.
Large language models (LLMs) show promise in assisting with surgical decisions but require further development for clinical reliability. Human surgeons consistently outperformed AI in accuracy and overall usefulness for maxillofacial trauma cases.
Area of Science:
- Medical Artificial Intelligence
- Surgical Decision Support Systems
- Maxillofacial Traumatology
Background:
- Large language models (LLMs) present potential benefits for healthcare professionals in diagnostic and therapeutic decision-making.
- However, their integration into clinical workflows raises concerns regarding usefulness, reliability, and ethical implications.
Purpose of the Study:
- To evaluate the capabilities of LLMs, specifically ChatGPT-4 and Google Bard, in managing complex surgical scenarios within maxillofacial traumatology.
- To compare the performance of LLMs against human maxillofacial surgery residents in real-world clinical case simulations.
Main Methods:
- A cross-sectional study involving 30 real-world maxillofacial trauma cases.
- Cases were presented to ChatGPT-4, Google Bard, and surgical residents.
- Performance was assessed by expert surgeons using the AIPI and QAMAI evaluation tools.
Main Results:
- ChatGPT-4 and Bard demonstrated similar abilities in considering patient features but varied in differential diagnosis suggestions.
- ChatGPT-4 showed superiority over Bard in proposing additional examinations and treatment plans.
- Maxillofacial surgery residents consistently achieved higher scores than LLMs across all QAMAI parameters, including accuracy, clarity, and overall usefulness.
Conclusions:
- LLMs show potential as supportive tools for clinical decision-making in facial traumatology.
- Significant advancements are needed to ensure the reliability of LLMs for practical clinical application.
- The AIPI and QAMAI tools are valuable for assessing LLM responses, emphasizing the need for standardized evaluation methods.
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Muscles for Facial Expressions

