Related Experiment Video
Updated: Jun 20, 2026

14:05
One Dimensional Turing-Like Handshake Test for Motor Intelligence
Published on: December 15, 2010
27.7K
Can AI-Based ChatGPT Models Accurately Analyze Hand-Wrist Radiographs? A Comparative Study
Ahmet Yıldırım1, Orhan Cicek1, Yavuz Selim Genç2
1Department of Orthodontics, Faculty of Dentistry, Zonguldak Bulent Ecevit University, Zonguldak 67600, Türkiye.
Diagnostics (Basel, Switzerland)
|June 26, 2025
Summary
Large language models (LLMs) show promise in predicting bone age and growth stages, offering a practical alternative to traditional methods. While not replacing clinical exams, these AI tools demonstrate potential for preliminary assessments in orthodontics.
Area of Science:
- Orthodontics and Dental Anthropology
- Artificial Intelligence in Medicine
- Radiographic Analysis
Background:
- Accurate bone age assessment is crucial for orthodontic treatment planning.
- Conventional methods and deep learning models require significant infrastructure and training.
- Large language models (LLMs) offer a potential infrastructure-independent alternative.
Purpose of the Study:
- To evaluate the effectiveness of LLM-based chatbot systems (ChatGPT) in predicting bone age and growth stages.
- To explore LLMs as practical, infrastructure-independent alternatives to current methods.
- To compare LLM performance against conventional methods and convolutional neural network (CNN) models.
Main Methods:
- Three ChatGPT models (GPT-4o, GPT-o4-mini-high, GPT-o1-pro) analyzed 90 anonymized hand-wrist radiographs.
- Radiographs represented pre-peak, peak, and post-peak growth stages, with equal sex distribution.
- Expert orthodontists established reference standards using Fishman's SMI system and Greulich-Pyle Atlas.
Main Results:
- All LLMs demonstrated significant agreement (p < 0.001) in bone age prediction; GPT-o1-pro showed highest concordance (r=0.546).
- GPT-o4-mini-high achieved 72.2% accuracy within a ±2 year deviation for bone age.
- GPT-4o exhibited the highest agreement (κ=0.283, p < 0.001) for growth stage classification.
Conclusions:
- General-purpose LLMs can assist in predicting bone age and growth stages, each with unique strengths.
- LLMs offer contextual reasoning and preliminary assessment capabilities without domain-specific training.
- Further development is needed, but LLMs show potential as supportive tools in orthodontics, complementing clinical examination.

