Related Experiment Video
Updated: Aug 6, 2026

09:10
Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Large Language Models for Traumatic Dental Injuries Across Web-Based and Mobile-Based Interfaces: Assessing Accuracy,
Ezgi Can Cekic1, Mertkan Kumru1, Burcu Yilmaz1
1Department of Endodontics, Faculty of Dentistry, Usak University, Usak, Türkiye.
Clinical and Experimental Dental Research
|July 19, 2026
Summary
Large language models (LLMs) can aid in accessing information on traumatic dental injuries (TDIs). Qwen2.5-Max showed strong performance across interfaces, but all LLM outputs require professional verification for clinical use.
Area of Science:
- Artificial Intelligence in Dentistry
- Clinical Decision Support Systems
- Information Retrieval for Traumatic Dental Injuries
Background:
- Traumatic dental injuries (TDIs) necessitate rapid, guideline-based clinical decisions.
- Accessing accurate information on TDIs can be challenging for practitioners.
- Large language models (LLMs) are emerging as quick online information resources.
Purpose of the Study:
- To evaluate the accuracy and consistency of multiple LLMs in answering TDI-related questions.
- To assess the influence of different user interfaces (web vs. mobile) on LLM performance.
- To compare the performance of ChatGPT-4o, DeepSeek-V3, Gemini 2.0 Flash, and Qwen2.5-Max.
Main Methods:
- Twenty TDI questions based on 2020 IADT guidelines were formulated (10 open-ended, 10 yes-no).
- Four LLMs were queried via web and mobile interfaces over five days, generating 800 responses.
- Open-ended answers were scored using GQS and mDISCERN; yes-no answers were compared to a key.
Main Results:
- Qwen2.5-Max achieved higher Global Quality Scores (GQS) and modified DISCERN (mDISCERN) scores.
- Yes-no question accuracy ranged from 86% to 91% with no significant model differences.
- ChatGPT-4o performed better on the web, while Qwen2.5-Max excelled on mobile; Qwen2.5-Max showed higher temporal consistency.
Conclusions:
- LLMs can be supplementary tools for guideline-based TDI information, particularly for closed-ended questions.
- LLM performance varies by model, interface, and question type.
- All LLM-generated information for TDIs must be cautiously interpreted and verified by dental professionals.
Keywords:
artificial intelligencedentistryendodonticsinformation reliabilitynatural language processingtraumatic dental injuries
