Related Experiment Video
Updated: Feb 4, 2026

Measurement of X-ray Beam Coherence along Multiple Directions Using 2-D Checkerboard Phase Grating
Published on: October 11, 2016
Intelligent identification of osteoporosis on hip X-rays using vision transformer
Wei Huang1,2, Pin Pan1,2, Kunpeng Liu3
1Hefei Hospital Affiliated to Anhui Medical University, Hefei, Anhui, 230011, China.
Objective:
This study aimed to develop and evaluate a deep learning model based on the Vision Transformer (ViT) architecture for the automatic classification of hip X-ray images into three categories: normal bone mass, osteopenia, and osteoporosis. The goal was to explore the model's potential for early screening and auxiliary diagnosis of osteoporosis.
Methods:
A total of 3016 hip anteroposterior X-ray images were retrospectively collected from Hefei Hospital Affiliated to Anhui Medical University and affiliated community clinics. After standard preprocessing and extraction of proximal femur regions of interest (ROI), the dataset was split into training and internal validation sets in an 8:2 ratio. A pretrained ViT model was fine-tuned for the three-class classification task and compared with conventional convolutional neural networks (ResNet50 and InceptionV3). Performance was assessed using accuracy, area under the ROC curve (AUC), sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Additionally, the model was further validated using an external dataset to assess its generalizability.
Results:
On the internal validation set, the ViT model achieved an overall classification accuracy of 97.0%. The AUCs for osteoporosis, osteopenia, and normal bone mass were 99.6%, 99.4%, and 99.9%, respectively. The PPV were 96.9%, 94.1% and 100%;The NPV were 97.9%, 98.5% and 99.2%. On the external validation set, the ViT model achieved an overall classification accuracy of 89.4%. The AUCs for osteoporosis, osteopenia, and normal bone mass were 96.5%, 91.6%, and 98.4%, respectively. The PPV were 83.5%, 90.2% and 91.3%;The NPV were 94.5%, 91.3% and 96.4%. The model demonstrated high sensitivity, specificity, PPV, and NPV across all classes, and outperformed both ResNet50 and InceptionV3 in overall diagnostic performance and classification stability.
Conclusion:
The ViT-based deep learning model showed excellent performance in classifying bone mineral density using hip X-rays, with high accuracy and generalizability. Relying on routine X-ray images, this method provides a cost-effective and efficient tool for osteoporosis screening, with strong potential for clinical implementation in primary care settings.
Related Concept Videos
X-ray Crystallography
Diffraction
Diffraction is the change in the direction of travel experienced by an electromagnetic wave when it encounters a physical barrier whose dimensions are comparable to those of the wavelength of the light. X-rays are electromagnetic radiation with wavelengths about as long as the distance between neighboring...
Vision
Intelligence
Color Vision
Bacterial Transformation
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...

