Related Experiment Videos
Development and performance of a specialized large language model for restorative treatment planning
1Assistant Professor, Department of Restorative Dentistry, Maurice H. Kornberg School of Dentistry, Temple University, Philadelphia, Pa.
Statement Of Problem:
Although general-purpose large language models (LLMs) have shown some promise in supporting dental treatment planning, their performance remains suboptimal. This limitation underscores the need for specialized models tailored to restorative dentistry.
Purpose:
The purpose of this study was to develop and evaluate a domain-specific LLM to provide evidence-based restorative treatment planning (RTP) for endodontically treated teeth (ETT).
Material And Methods:
The pipeline included knowledge collection from a dataset based on 39 clinical factors, textbooks, and peer-reviewed literature to establish RTP-GPT; supervised fine-tuning with clinical treatments; integration of a retrieval-augmented generation (RAG) system; and benchmarking against general-purpose LLMs. Twenty novel scenarios were tested across 4 models. Two blind evaluators scored outputs in 5 domains: accuracy, conservativeness, recommendation completeness, justification completeness, and analytical reasoning. The Kruskal-Wallis and Dunn post hoc tests assessed differences, and intraclass correlation coefficients (ICCs) measured reproducibility across 3 testing days (α=.05).
Results:
RTP-GPT achieved the best overall performance, significantly outperforming ChatGPT and RTP-GPT+RAG in all domains (Bonferroni-corrected P<.05). Gemini ranked second overall and showed stronger analytical reasoning than RTP-GPT+RAG but was less conservative than RTP-GPT (P<.001). RTP-GPT+RAG surpassed ChatGPT in conservativeness and the completeness of recommendations and justifications. Regarding reliability, RTP-GPT+RAG demonstrated the highest reproducibility with good-to-excellent ICCs, while RTP-GPT, despite superior accuracy, showed the lowest reproducibility in accuracy, conservativeness, and treatment completeness.
Conclusions:
RTP-GPT provided the strongest performance in RTP for ETT, outperforming Gemini in conservativeness and both RTP-GPT+RAG and ChatGPT across all domains. However, its limited reproducibility underscores the need for further considerations.