Related Experiment Video
Updated: Oct 7, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Device-Specific Conformal Reliability under Scanner-Level Acquisition Shift in Breast Ultrasound CAD
1Department of Actuarial Sciences, Selcuk Universitesi Fen Fakultesi, Selcuk University, Faculty of Science, Konya, Konya, 42130, Turkey.
Abstract:
Conformal prediction provides finite-sample coverage control under exchangeability, but empirical reliability may change across scanner-defined deployment domains. We audit three conformal procedures and one non-conformal abstention baseline on leave-one-device-out (LODO) partitions of BUS-BRA using EfficientNet-B0, ResNet-50 and ConvNeXt-Tiny. All conformal thresholds are computed by direct finite-sample order-statistic indexing. The primary prediction and score unit is the image; source splits are patient-disjoint, target calibration/test partitions are patient-disjoint within each recalibration draw, and patient-cluster bootstrap plus patient-aggregated analyses assess within-patient dependence. At nominal α=0.05, EfficientNet-B0 Mondrian CP gives malignant set-miss rates of 0.004 for GE Logiq 7, 0.002 for GE Logiq 5 and 0.244 for Toshiba; the Toshiba patient-cluster 95% interval is 0.140-0.359. Patient aggregation preserves the contrast (0.194 for Toshiba versus 0.000 for both GE devices). In paired target recalibration, k=10 lowers Toshiba mean miss rate from 0.246 to 0.057 while increasing abstention to 0.603, whereas the same intervention raises the GE means to 0.060 and 0.067. Across the 500 Toshiba k=10 seed-draw combinations, the post-recalibration median miss rate is 0.021 (IQR 0.000-0.063), and 96.4% of draws improve relative to their paired pre-recalibration value. A finite-resolution analysis shows that α=0.02 is rank-unavailable for all tested few-shot conditions; under direct finite-sample order-statistic indexing, α=0.02 and α=0.05 coincide at k=5 and k=10 because both select the observed class maxima, but they separate in the larger k=20 condition. These results document scanner-level differences in empirical conformal reliability and motivate evaluating target recalibration when an independent target-domain reliability audit provides evidence of undercoverage, rather than establishing a prospective trigger rule.

