Related Experiment Video
Updated: Feb 14, 2026

Evaluation of Commercial-Off-The-Shelf Wrist Wearables to Estimate Stress on Students
Published on: June 16, 2018
Atlas-Assisted Bone Age Estimation from Hand-Wrist Radiographs Using Multimodal Large Language Models: A Comparative
1Department of Radiology, Kastamonu Training and Research Hospital, Kastamonu 37150, Turkey.
ChatGPT-5 demonstrated the best performance among four multimodal large language models (LLMs) in bone age assessment, showing clinical utility. Other models like Gemini 2.5 Pro, Grok-3, and Claude 4 Sonnet showed insufficient reliability for bone age determination.
Area of Science:
- Medical Imaging
- Artificial Intelligence in Medicine
- Pediatric Endocrinology
Background:
- Bone age assessment is crucial in pediatric endocrinology and forensic medicine.
- Multimodal large language models (LLMs) show promise in medical imaging but require evaluation for bone age determination.
- The Gilsanz-Ratib (GR) atlas is a standard reference for bone age assessment.
Purpose of the Study:
- To evaluate the diagnostic performance of four multimodal LLMs (ChatGPT-5, Gemini 2.5 Pro, Grok-3, Claude 4 Sonnet) in bone age determination.
- To compare LLM performance against a radiologist's assessment using the GR atlas.
- To determine the clinical utility of these LLMs for bone age assessment.
Main Methods:
- Retrospective analysis of 245 pediatric wrist radiographs (age < 18).
- LLMs used GR atlas-assisted prompting for bone age estimation.
- Performance metrics included Mean Absolute Error (MAE), Intraclass Correlation Coefficient (ICC), and Bland-Altman analysis.
- Comparison with an experienced radiologist's GR atlas-based assessment.
Main Results:
- ChatGPT-5 achieved the lowest MAE (1.46 years) and highest ICC (0.849), demonstrating superior performance.
- Gemini 2.5 Pro showed moderate performance (MAE: 2.24 years).
- Grok-3 (MAE: 3.14 years) and Claude 4 Sonnet (MAE: 4.29 years) exhibited error rates too high for clinical application.
Conclusions:
- Significant performance variations exist among multimodal LLMs for bone age assessment.
- Only ChatGPT-5 met criteria for clinical usefulness, suggesting potential as an auxiliary tool or for educational support under supervision.
- Current multimodal LLMs, excluding ChatGPT-5, lack the reliability for clinical bone age determination.
More Related Videos
07:22Glycemic Impact on Knee Osteoarthritis Symptoms on Physical, Radiographic, and Inflammatory Markers among Individuals Aged 50 and Over with Diabetes
Published on: March 7, 2025
04:10Author Spotlight: Expanding Interventional Pulmonology Research with Robotic-Assisted Bronchoscopy
Published on: July 19, 2024
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Hand hygiene
Hand washing...
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...