Related Experiment Video
Updated: Sep 27, 2026

Measuring Maxillary Posterior Tooth Movement: A Model Assessment using Palatal and Dental Superimposition
Published on: February 23, 2024
Classification Performance of General-Purpose Multimodal Large Language Models Across Orthodontic Radiographic Tasks:
Nuri Can Tanrısever1, Kutluhan Yılmaz1, Ayşegül Dilara Güvenç Tokur1
1Department of Orthodontics, School of Dentistry, Ankara University, 06560 Ankara, Türkiye.
Abstract:
Background and Objectives: General-purpose multimodal large language models (MLLMs) can interpret radiographic images, but their classification performance across orthodontic tasks remains uncertain. This study compared the classification performance of ChatGPT, Gemini, and Claude on lateral cephalometric, hand-wrist, and panoramic radiographs. Materials and Methods: This retrospective diagnostic accuracy study included 250 individuals, each contributing one lateral cephalometric, hand-wrist, and panoramic pretreatment radiograph (750 total). Reference classifications were established by two experienced orthodontists, with disagreements resolved by consensus. Lateral cephalometric radiographs were classified as skeletal Class I, II, or III based on the ANB angle according to Steiner analysis; hand-wrist radiographs as prepubertal, pubertal, or postpubertal; and panoramic radiographs as early mixed, late mixed, or permanent dentition. Each image was evaluated once by each AI platform using identical Turkish prompts in separate chat sessions. Classification accuracy, balanced accuracy, macro-F1, class-specific metrics, and reference agreement were assessed. Generalized estimating equations (GEE) assessed platform, radiograph type, and interaction effects on correct classification. Results: ChatGPT had the highest hand-wrist accuracy (81.6%; 95% CI, 76.3-85.9), whereas Gemini had the highest panoramic accuracy (92.8%; 95% CI, 88.9-95.4). Lateral cephalometric accuracies were 70.0%, 64.4%, and 64.8% for ChatGPT, Gemini, and Claude, respectively, with no significant interplatform difference (p = 0.336). The platform × radiograph type interaction was significant (Wald χ2 = 42.52; df = 4; p < 0.001). Agreement with the reference standard was highest for ChatGPT on hand-wrist radiographs (κw = 0.758) and Gemini on panoramic radiographs (κw = 0.883). Conclusions: Classification performance was task- and platform-dependent, with no model consistently achieving the highest performance. For the predefined classification tasks, these models should not be used as standalone tools for these classification tasks. Their potential as decision-support tools requires prospective evaluation of AI-assisted clinician performance and external validation.