Related Experiment Video
Updated: Jan 14, 2026

A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
Evaluating the Effectiveness of Large Language Models in Addressing Patient Queries Regarding Maxillomandibular
Ragavi Alagarsamy1, Babu Lal2, Jitendra Chawla3
1Project Research Scientist-1, Department of Oral and Maxillofacial Surgery, Maulana Azad Institute of Dental Sciences, New Delhi, India.
Background:
Patients with maxillofacial fractures increasingly seek information from large language models (LLMs), yet the accuracy and readability of these responses remain uncertain.
Purpose:
This study evaluated the performance of 5 publicly accessible LLMs in answering frequently asked questions (FAQs) about maxillomandibular fixation (MMF).
Study Design, Setting, And Sample:
This in-silico cross-sectional study, conducted in January 2025, evaluated 47 FAQs and yielded 235 responses from 5 open-access LLMs, excluding subscription-based models.
Predictor Variable:
The predictor variable was LLM architecture: decoder-only transformer models (DOT-1, DOT-2), a multimodal transformer model (MTM), a productivity-focused model (PM), and a constitutional artificial intelligence (AI)-based model (CAM).
Outcome Variables:
The primary outcome was LLM performance, measured with the QUEST (Quality of information, Understanding and reasoning, Expression style and persona, Safety and harm, and Trust and confidence) framework. Domains assessed were accuracy (Likert ≥4), hallucination (presence/absence of fabricated content), usefulness, clarity, trust, and satisfaction (Likert 1 to 5), and readability (Flesch-Kincaid Reading Ease [FKRE] and Grade Level [FKGL]). Responses were rated independently by 7 evaluators (5 oral and maxillofacial surgeons and 2 residents) in a blinded manner.
Covariates:
None.
Analyses:
Ordinal outcomes were analyzed with the Friedman test and pairwise Wilcoxon signed-rank tests. Readability was compared with one-way ANOVA. Inter-rater reliability was measured with Fleiss' kappa. Statistical significance was set at P < .05.
Results:
The sample included 235 LLM-generated responses. DOT-1 showed the highest accuracy (88.5 ± 6.2%), which was statistically significantly greater than DOT-2 (79.6 ± 10.1%) and PM (81.2 ± 9.3%) (P = .004). It also had a statistically significantly lower hallucination rate (5.2%) compared with DOT-2 (10.1%) and PM (9.4%) (P = .013). CAM performed comparably in accuracy (86.3 ± 7.1%); however, its readability was statistically significantly poorer (Flesch-Kincaid Grade Level = 22.7 ± 12.9; P < .001). Multimodal transformer model showed intermediate performance. Inter-rater agreement was almost perfect for accuracy (κ = 0.79 to 1.00) and hallucination (κ = 0.91 to 1.00) and moderate to substantial for ordinal variables.
Conclusion And Relevance:
LLMs can provide accurate responses to maxillomandibular fixation queries, but readability remains limited and model-dependent. These findings underscore the need for developing more patient-friendly artificial intelligence (AI) outputs and highlight the importance of clinician oversight in guiding patients' use of LLMs.

