Related Experiment Video
Updated: Mar 21, 2026

Author Spotlight: Accelerating Cognitive Impairment Research in Mice Through Stereotaxic Injection and Cost-Effective Dyes
Published on: July 19, 2024
Expert variability as a benchmark for validating automatic CT-MRI registration in brain stereotactic radiosurgery
Valeria Faccenda1, Denis Panizza1, Valentina Pinzi2
1Medical Physics, Fondazione IRCCS San Gerardo Dei Tintori, Monza, Italy; School of Medicine and Surgery, University of Milan Bicocca, Milan, Italy.
Background:
Accurate CT-MRI registration is critical for stereotactic radiosurgery (SRS) of brain metastases (BM), but even expert manual fusion shows variability. Without a true ground truth, algorithm validation requires a clinically meaningful benchmark. This study aimed to define a probabilistic gold standard (GS) for rigid registration and to establish acceptance thresholds for automatic algorithms.
Methods:
Twenty CT-MRI pairs (39 BM) were registered twice by six operators (n = 240). Variability was assessed by bootstrap resampling to estimate 99% confidence intervals (CIs) for rotational and translational mean absolute error (MAE), BM barycenter shift, DICE, and HD. For each case, the GS was defined as the median translational and rotational components across experts. Deviations from the GS were bootstrapped to derive expert-to-GS variability, used as acceptance range. Three mutual-information (MI)-based methods were tested: standard MI, skull-box MI, and contour-based (CB) MI (manual or AI-generated contours). Less-experienced operators were also evaluated.
Results:
Experts showed low but non-negligible variability (99% CI: 0.31° rotation, 0.34 mm translation; max: 0.92°, 1.19 mm; median barycenter shift: 0.7 mm, range 0.0-2.8 mm). Only the CB algorithm achieved median values within expert acceptance for all metrics except HD, with no significant differences between manual and AI-generated contours (P > 0.05). Less-experts showed larger deviations, markedly reduced when starting from CB-aligned datasets.
Conclusions:
Quantifying multi-expert variability enables a probabilistic GS and clinically relevant thresholds for CT-MRI registration. This benchmark reflects human performance and provides a framework for algorithm validation. An AI-driven CB algorithm demonstrated high reliability, reduced operator dependence, and may streamline SRS workflows.
More Related Videos
10:25Brain Infarct Segmentation and Registration on MRI or CT for Lesion-symptom Mapping
Published on: September 25, 2019
07:57Positron Emission Tomography-based Dose Painting Radiation Therapy in a Glioblastoma Rat Model using the Small Animal Radiation Research Platform
Published on: March 24, 2022