Related Experiment Video
Updated: Jul 4, 2026

A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
[Recommendation and comparison of four large language models for prosthetic treatment options in edentulous patients]
1Department of Orthodontics, School of Stomatology, The Fourth Military Medical University, National Clinical Research Center for Oral Diseases, State Key Laboratory of Oral Maxillofacial Reconstruction and Regeneration, Shaanxi Key Laboratory of Stomatology, Xi'an 710032, China.
Abstract:
Objective: To evaluate the comprehensiveness and consistency of treatment plan recommendations provided by four large language models (LLM) for edentulous patients, and to explore their feasibility and limitations as clinical decision-making support tools. Methods: Twelve standardized simulated clinical records of edentulous patients were constructed, each fully incorporating information on six dimensions influencing treatment plan selection (local oral conditions, general health status, socioeconomic factors, patient preferences, convenience of maintenance, and long-term maintenance capability). Four LLM (ChatGPT, Tongyi Qianwen, Wenxin Yiyan, DeepSeek) were used to provide treatment recommendations and corresponding rationales for implant-supported prostheses and complete dentures (single-turn question-and-answer, fixed prompts, each LLM responded three times to each of the 12 standardized records, with the conversation history cleared before each response). Meanwhile, a panel of three chief physicians specializing in prosthodontics and implant dentistry was invited to provide expert opinions for each scenario. Subsequently, two senior attending physicians independently evaluated the responses of each LLM using a multidimensional scoring system (consistency of recommendations, completeness of rationales, risk identification capability, and decision-making transparency). Finally, the Kruskal-Wallis test was used to analyze differences among the LLM. Results: The inter-rater reliability between the two senior attending physicians was excellent (Spearman's ρ=0.906, P0.001). The performance of LLM varied across dimensions. ChatGPT achieved the highest overall score of 3.13 (2.75, 3.44), followed by DeepSeek 2.75 (2.50, 2.94), Tongyi Qianwen 2.38 (2.00, 2.50), and Wenxin Yiyan 2.13 (1.56, 2.43). The overall difference among the four models was statistically significant (H=22.80, P0.001). Regarding consistency of recommendations, ChatGPT showed the highest concordance with expert opinions (the combined proportion of basic and complete consistency was 12/12). In terms of completeness of rationales, the mention rates of local oral conditions, general health status, and socioeconomic factors were all ≥6/12 for each model (with ChatGPT achieving ≥11/12). However, for convenience of maintenance and long-term maintenance capability, the mention rates for Wenxin Yiyan, Tongyi Qianwen, and DeepSeek were ≤6/12. Conclusions: The four LLM showed significant differences in their performance in edentulous restoration decision-making. ChatGPT achieved the highest overall score and the best concordance with expert opinions. Regarding non-medical factors such as convenience of maintenance and long-term maintenance capability, the mention rates of Wenxin Yiyan, Tongyi Qianwen, and DeepSeek were relatively low. These findings suggest that current LLM are not yet capable of comprehensively weighing all key decision-making factors, with a notable deficiency in assessing non-medical factors in particular. They cannot replace the judgment of clinical practitioners; however, they demonstrate certain supportive value in structured information presentation and consistency of recommendations.
