Related Experiment Video
Updated: Oct 29, 2025

Software-Assisted Quantitative Measurement of Osteoarthritic Subchondral Bone Thickness
Published on: March 18, 2022
Interrater and Intrarater Reliability of the Vertebral Bone Quality Score
Andrew T Schilling1, Jeff Ehresman2, Zach Pennington3
1Department of Neurosurgery, Johns Hopkins University School of Medicine, Baltimore, Maryland, USA.
This study evaluated how consistently different medical professionals and trainees could calculate the Vertebral Bone Quality (VBQ) score using magnetic resonance imaging. The researchers found that the scores were highly reliable regardless of the rater's experience level, suggesting this method is a stable tool for assessing bone health without radiation.
Area of Science:
- Orthopedic surgery and Vertebral Bone Quality score research within musculoskeletal diagnostics
- Radiological imaging and clinical reliability assessment
Background:
Clinicians often struggle to accurately predict patient outcomes following spinal fusion procedures due to limited bone density assessment tools. Conventional diagnostic methods frequently expose patients to ionizing radiation during routine preoperative screening. That uncertainty drove interest in non-invasive magnetic resonance imaging techniques for evaluating skeletal integrity. The Vertebral Bone Quality score emerged as a potential solution for quantifying bone health in clinical settings. However, the consistency of this metric across diverse medical personnel remained poorly characterized in existing literature. No prior work had resolved whether trainees could achieve the same precision as experienced specialists when applying this scoring system. This gap motivated a systematic investigation into the reproducibility of these measurements. Understanding the stability of such metrics is a prerequisite for their widespread adoption in surgical planning.
Purpose Of The Study:
This investigation aimed to determine the intrarater and interrater reliability of the Vertebral Bone Quality score among medical professionals and trainees. The researchers sought to establish whether this magnetic resonance imaging-based metric could be applied consistently by individuals with varying levels of clinical experience. Spinal fusion surgery outcomes often depend on preoperative bone health, yet existing assessment methods frequently involve significant radiation exposure. This study addressed the need for a reliable, non-invasive alternative to traditional diagnostic testing. By evaluating the consistency of the scoring system, the authors intended to validate its potential for widespread clinical implementation. The team hypothesized that the metric would demonstrate high reproducibility despite differences in rater background. This work provides necessary evidence regarding the stability of the scoring process in a clinical environment. Ultimately, the study serves to clarify whether this tool can be reliably utilized across different levels of medical training.
Main Methods:
The research team recruited thirteen reviewers with varying levels of clinical expertise to participate in the study. These individuals performed assessments on thirty distinct patient imaging volumes at two separate time points. A two-month interval separated these sessions to minimize recall bias during the evaluation process. The study design employed two-way random effects modeling to calculate the intraclass correlation coefficient for all participants. Square-weight Cohen kappa and Kendall Tau-b tests served to verify the consistency of scoring assignments across both rounds. The investigators included patient volumes representing both degenerative and oncologic conditions to ensure broad applicability. This approach allowed for a robust comparison of interrater and intrarater performance among the diverse group of raters. The methodology focused on quantifying the stability of the scoring system rather than its diagnostic accuracy.
Main Results:
The primary analysis revealed that all participants achieved moderate to excellent reliability for the total score. Intraclass correlation coefficients for individual raters ranged from 0.667 to 0.957, while Cohen kappa values spanned 0.648 to 0.921. The constituent components of the score demonstrated excellent reliability, with all intraclass correlation coefficient values reaching at least 0.97. Interrater agreement remained consistently high during the first round of assessment, yielding an intraclass correlation coefficient of 0.818. During the second round, the interrater reliability remained stable with an intraclass correlation coefficient of 0.800. The data indicated no significant correlation between the level of rater training and the reliability of the scores. These results highlight the robust nature of the metric across different observers. The findings confirm that the scoring system maintains high reproducibility regardless of the rater's professional background.
Conclusions:
The Vertebral Bone Quality score demonstrates strong consistency when applied by both expert clinicians and medical trainees. These findings suggest that the metric provides a stable, reproducible assessment of skeletal status across different observers. The high intraclass correlation coefficients indicate that individual raters maintain their scoring patterns over time. Furthermore, the lack of association between experience levels and reliability supports the broad utility of this diagnostic tool. Future research must now confirm these results through external validation in larger, more diverse patient cohorts. Investigators should also prioritize studies linking these scores to actual biomechanical properties of the spine. Such efforts will clarify the clinical significance of the metric in predicting surgical success. This work establishes a foundation for integrating non-invasive bone quality assessment into routine preoperative workflows.
Frequently Asked Questions
The researchers utilized intraclass correlation coefficients and Cohen kappa statistics to quantify agreement. They observed that individual raters achieved scores ranging from 0.667 to 0.957, while group consistency remained high at 0.818 and 0.800 during the two assessment rounds.
The team analyzed magnetic resonance imaging volumes from thirteen participants. These individuals represented various medical specialties and training stages, ranging from trainees to experienced professionals, to ensure a comprehensive evaluation of the scoring process.
The study required two separate evaluation sessions spaced two months apart. This temporal separation allowed the researchers to effectively measure intrarater stability, ensuring that the same rater produced consistent results over an extended period.
The investigators processed imaging data from thirty patients. These subjects presented with either degenerative spinal conditions or oncologic indications, providing a diverse clinical dataset for testing the robustness of the scoring methodology.
The authors measured the intraclass correlation coefficient for all constituent components of the score. They reported excellent reliability for these individual parts, with values consistently reaching or exceeding 0.97 across all raters.
The researchers propose that this scoring method offers a reliable alternative to radiation-heavy testing. They suggest that future studies should focus on validating these findings externally and determining how well the score models actual bone biomechanical properties.
Related Concept Videos
Classification of Bones
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT

