Related Experiment Video
Updated: Jun 27, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Feasibility of a multi-metric framework for evaluating patient-facing AI communication in cosmetic dentistry: an
Alaa Al-Haddad1, Omar Al-Karadsheh2, Yazan Hassona2
1Department of Restorative Dentistry, School of Dentistry, The University of Jordan, Amman, Jordan.
Background:
Large language models (LLMs) are increasingly used by the public to obtain oral health information, yet reproducible methods to benchmark the communication quality of patient-facing outputs remain underdeveloped. Prior evaluations have focused mainly on factual accuracy and guideline concordance, while giving less attention to whether responses are understandable, actionable, empathetic, well structured, and bounded by appropriate safety messaging. This gap is especially relevant in cosmetic dentistry, where patients often make elective and potentially irreversible decisions based on online information.
Methods:
This proof-of-concept comparative benchmarking study used a consolidated 80-prompt test set derived from thematic analysis of real-world patient inquiries and cross-LLM synthesis across four cosmetic dentistry domains: tooth whitening, veneers, implants, and orthodontic aligners. Responses from an instruction-configured assistant (CSA-GPT) and a general-purpose baseline (ChatGPT5.2) were generated under controlled conditions, yielding 160 responses. Two board-certified specialists independently evaluated all responses using a theory-informed, exploratory 20-point rubric assessing readability (Flesch-Kincaid Grade Level, FKGL), ethical disclaimer inclusion, practicality, empathetic tone, and structural clarity. A separate clinical safety audit assessed major factual errors and critical omissions. Between-model comparisons used paired analyses with effect sizes, and linear mixed-effects models examined Model, Domain, and Model × Domain interaction.
Results:
CSA-GPT outperformed ChatGPT5.2 across all evaluated communication metrics. Mean total rubric score was 17.95 ± 1.62 for CSA-GPT vs. 9.55 ± 1.94 for ChatGPT5.2 (p < 0.001; Cohen's d = 3.22). Mean FKGL was lower for CSA-GPT (6.07 ± 1.28) than for ChatGPT5.2 (9.12 ± 1.71; p < 0.001). Practicality, empathetic tone, and structural clarity were all significantly higher for CSA-GPT (all p < 0.001). Mandatory disclaimers were present in 100% of CSA-GPT responses and 0% of ChatGPT5.2 responses. Safety audit error rates were low and did not differ significantly between models (1.25% vs. 3.75%, p = 0.25). Mixed-effects models confirmed a strong overall model effect, with domain-dependent interactions for readability, empathy, and structural clarity.
Conclusions:
In this exploratory proof-of-concept study, instruction configuration improved the patient-facing communication quality of LLM responses in cosmetic dentistry across readability, practicality, empathetic tone, structural clarity, and safety boundary-setting, without increasing major factual errors. These findings support the feasibility of a multi-metric benchmarking approach for evaluating patient-facing dental AI, while highlighting the need for psychometric refinement, external validation, and broader testing before such approaches can inform governance or implementation.
