Related Experiment Video
Updated: Aug 28, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Patient-Facing AI Chatbot Treatment-Direction Advice in Orthodontic Health Communication: A Scenario-Based Comparison
Neslihan Karaoğlan1, Hakan Karaoğlan2
1Dental Clinic, University of Health Sciences, Sultan II. Abdülhamid Han Training and Research Hospital, 34668 Istanbul, Türkiye.
Abstract:
Background/Objectives: AI chatbots may shape patient expectations before professional consultation. This scenario-based first-response study evaluated whether four user-facing chatbots provided orthodontic treatment-direction advice concordant with an expert benchmark and whether responses contained safety, referral, or overconfidence concerns. Methods: Forty fictional Turkish patient-oriented scenarios across eight categories were independently coded by three orthodontists as clear aligners, fixed appliances, both options, examination required, or advanced specialist/surgical evaluation required. Each scenario was submitted once to ChatGPT, Claude, Copilot, and Gemini on 20 May 2026. Two independent non-author orthodontists coded 160 archived first responses using a predefined framework, with adjudication before analysis. Results: Inter-expert agreement was moderate (Fleiss kappa = 0.491; Gwet AC1 = 0.528). Under the majority benchmark, exact concordance was 82.5% for ChatGPT, 67.5% for Claude, 42.5% for Copilot, and 37.5% for Gemini (Cochran Q = 34.105, p < 0.001). The overall difference remained significant in the 17 unanimous scenarios (Q = 11.455, p = 0.010), but a post hoc alternative-reference analysis that adopted the dissenting expert code in the 23 non-unanimous scenarios attenuated the rates to 57.5%, 52.5%, 52.5%, and 42.5%, respectively (Q = 4.222, p = 0.238). Coded safety-concern rates ranged from 15.0% to 62.5%. Conclusions: The sampled first responses differed in treatment direction and safety coding, but estimates were sensitive to the expert reference definition. Under the tested single-date, single-language, and single-run conditions, the findings represent a conditional snapshot rather than a time-invariant ranking of model capability. Patient-facing chatbots should support nondirective pre-consultation education and referral, not autonomous appliance selection.
