Related Experiment Video
Updated: Oct 5, 2025

Cortical Bone Assessment Using Ultrasonic Guided Waves: A Reproducibility Study in a Healthy Population
Published on: January 31, 2025
External validation of deep learning-based bone-age software: a preliminary study with real world data
Winnah Wu-In Lea1, Suk-Joo Hong2, Hyo-Kyoung Nam3
1Department of Radiology, Guro Hospital, Korea University College of Medicine, Seoul, Republic of Korea.
This study tested how well a commercial computer program using artificial intelligence estimates bone age in children compared to human experts. Researchers found that while the software and human doctors sometimes disagreed, the computer's estimates were often as accurate or better than human experts when compared to the child's actual age.
Area of Science:
- Pediatric radiology outcomes research within deep learning-based bone-age assessment
- Medical informatics and diagnostic imaging validation studies
Background:
Clinicians frequently struggle with the time-consuming and complex nature of manual skeletal maturity evaluation. No prior work had resolved whether automated systems maintain accuracy when applied to diverse, real-world clinical patient populations. That uncertainty drove the need for rigorous external testing of commercial diagnostic tools. Prior research has shown that human interpretation of radiographic atlases often involves subjective variability among different specialists. This gap motivated an investigation into how modern computational models perform outside of controlled training environments. Previous studies often relied on curated datasets that may not reflect the messiness of daily hospital workflows. Such limitations hinder the widespread adoption of automated diagnostic support in pediatric healthcare settings. Researchers must therefore verify if these algorithms provide reliable results when facing the inherent noise found in standard clinical practice.
Purpose Of The Study:
The aim of this research was to evaluate the clinical performance of a commercially available deep learning-based software for skeletal maturity assessment. This study addressed the need to validate automated tools using real-world patient data rather than idealized datasets. The researchers sought to determine if such technology provides reliable results when applied to standard hospital workflows. By comparing machine output to human expert assessments, the team investigated the potential for clinical integration. This project specifically targeted the discrepancy between automated predictions and manual interpretations based on the traditional Greulich-Pyle atlas. The authors intended to clarify whether software can match the accuracy of specialists in a pediatric setting. This investigation was motivated by the desire to streamline time-intensive diagnostic tasks in radiology departments. The study provides evidence regarding the feasibility of deploying artificial intelligence to assist clinicians in daily practice.
Main Methods:
The review approach involved a retrospective analysis of 474 pediatric patients enrolled between late 2018 and early 2019. Investigators compared automated output against manual estimates provided by three specialists using the standard Greulich-Pyle atlas. Statistical verification utilized paired t-tests to identify differences between the machine and human observers. Researchers calculated Pearson's correlation coefficients to determine the strength of the relationship between various assessment methods. Bland-Altman plots served to visualize the agreement and potential bias between the software and the human experts. The team employed the intraclass correlation coefficient to quantify the degree of inter-rater reliability among the three human reviewers. Mean absolute error and root mean square error provided quantitative benchmarks for measuring accuracy against chronological age. This design ensured a comprehensive evaluation of the software's performance within a standard clinical environment.
Main Results:
Key findings from the literature indicate that the software achieved a strong correlation with human reviewers, reaching an r-value of 0.983. The root mean square error values between the software and the three human experts were 10.09, 10.76, and 13.06 months respectively. When compared to actual chronological age, the software yielded a root mean square error of 13.54 months. In contrast, the three human reviewers produced root mean square error values of 15.18, 16.19, and 19.53 months respectively. Bland-Altman analysis revealed a consistent tendency for both the software and human experts to overestimate biological age. The overall intraclass correlation coefficient among all human reviewers reached a value of 0.93. Statistical testing confirmed significant differences between the machine and human outputs with p-values below 0.025. The software demonstrated performance that was statistically similar or superior to human reviewers when evaluated against the actual age of the patients.
Conclusions:
The authors propose that their automated tool performs with a level of precision comparable to experienced human specialists. Synthesis and implications suggest that the software provides a viable alternative for routine skeletal maturity screening in busy clinics. Findings indicate that the program and human reviewers share a common tendency to slightly overestimate biological development. The researchers conclude that the model demonstrates robust reliability when measured against the actual chronological age of the pediatric subjects. This evidence supports the integration of such technology to assist practitioners in managing high patient volumes. The study highlights that the software maintains high correlation with expert human assessments despite statistically significant differences in absolute values. These results imply that automated systems can effectively augment clinical workflows without sacrificing diagnostic quality. Future implementation should consider these performance metrics when deploying digital health solutions in pediatric endocrinology departments.
Frequently Asked Questions
The researchers propose that the software achieves a high correlation (r=0.983) with human reviewers. While statistically significant differences exist between the computer and experts, the model demonstrates performance similar to or better than human clinicians when compared against the actual chronological age of the children.
The study utilized a commercially available deep learning-based software package. This tool was specifically designed to automate the assessment of skeletal development by analyzing radiographic images, which were then compared against the traditional Greulich-Pyle atlas standards used by the three human reviewers.
The researchers required three distinct human reviewers to provide an independent baseline for comparison. These included a musculoskeletal radiologist, a radiology resident, and a pediatric endocrinologist, ensuring a diverse range of clinical expertise to validate the software's output against standard practice.
The team employed real-world data collected from 474 children between November 2018 and February 2019. This data type is essential for assessing how algorithms handle the variability and noise inherent in standard clinical imaging workflows compared to idealized, curated training sets.
The researchers measured performance using the mean absolute error and root mean square error. These metrics quantified the discrepancy in months between the software's output and the human experts, as well as the deviation from the actual age of the pediatric patients.
The authors propose that their findings support the use of automated systems to assist in high-volume clinical settings. They suggest that such technology can effectively augment human decision-making, potentially reducing the time burden associated with manual skeletal maturity assessments in pediatric clinics.
More Related Videos
07:12Semiautomated Longitudinal Microcomputed Tomography-based Quantitative Structural Analysis of a Nude Rat Osteoporosis-related Vertebral Fracture Model
Published on: September 28, 2017
07:56Scanning Skeletal Remains for Bone Mineral Density in Forensic Contexts
Published on: January 29, 2018
Related Concept Videos
Classification of Bones
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
Changes in the Appendicular Skeleton with Age
Initially, the limb buds consist of a core of mesenchyme covered by a layer of ectoderm. The ectoderm at the end of the limb bud thickens to form a narrow crest called the apical ectodermal ridge. This ridge stimulates the underlying...