Related Experiment Video
Updated: May 26, 2026

Measuring Maxillary Posterior Tooth Movement: A Model Assessment using Palatal and Dental Superimposition
Published on: February 23, 2024
Accuracy and readability of artificial intelligence models in providing orthodontic retention related information: A
Afnan Ben Gassem1, Nebras Althagafi1
1Department of Preventive Dental Sciences, College of Dentistry, Taibah University, Almadinah Almunawara, Saudi Arabia.
Introduction:
Large language models (LLMs) are increasingly being used by patients for health information, yet their reliability in orthodontics remains uncertain. This study aims to evaluate the accuracy, reliability, quality, and readability of orthodontic retention information generated by ChatGPT 3.5, ChatGPT 4, Gemini, and Copilot.
Materials And Methods:
Twenty-three frequently asked questions about orthodontic retainers were collected and categorised into general retainer questions (n=8), fixed retainer questions (n=5), and removable retainer questions (n=10). Questions were entered into each AI model once under standardised conditions. Responses were anonymous and independently assessed by two consultant orthodontists. Accuracy was scored using a five-point Likert scale, reliability with the modified DISCERN tool, quality with the Global Quality Scale (GQS), and readability with the Flesch Reading Ease Score (FRES). Statistical analysis included ANOVA, Kruskal-Wallis, post-hoc tests, and intraclass correlation coefficients (ICC).
Results:
Evaluator agreement was excellent across all domains (ICC 0.821-0.957). ChatGPT 3.5 achieved the highest accuracy (mean 4.49), while ChatGPT 4 and Copilot scored highest in reliability (means 30.47 and 30.11). ChatGPT models outperformed Gemini and Copilot in quality, with over 75% of their responses rated good to excellent. Readability was low across all models; however, Copilot produced relatively more readable text (mean FRES score of 53.93).
Limitations:
This study is limited by its focus on single-turn responses which may not reflect the iterative interactions typical of real patient - AI conversations. In addition, the evolving nature of AI models may affect reproducibility, and its restriction to English, may limit its generalizability across languages.
Conclusion:
All AI models demonstrated moderate competence in providing orthodontic retention information, but their reliability was inconsistent, and readability was poor, necessitating human oversight and methodological refinement rather than serving as replacements for professional advice.

