Related Experiment Video
Updated: Jun 12, 2026

05:28
Clinical Imaging of Microwave Mammography
Published on: November 14, 2025
Side-level versus patient-level evaluation in four-view mammography classification: a comprehensive benchmark on the
1Department of Mathematics and Statistics, Qatar University, Qatar.
Future Science OA
|June 11, 2026
Summary
Evaluation granularity significantly impacts deep learning performance in mammography classification, with side-level metrics showing higher AUC than patient-level metrics. This highlights the need for standardized reporting in mammography research.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Biomedical Informatics
Background:
- Deep learning models show high performance in mammographic image classification.
- Inconsistent evaluation methods (side-level vs. patient-level) hinder reliable cross-study comparisons.
- The impact of evaluation granularity on reported performance needs quantification.
Purpose of the Study:
- To quantify the effect of evaluation granularity on mammography classification performance.
- To compare model architecture and evaluation methodology impacts on a single dataset.
- To establish uniform training and evaluation protocols for benchmarking.
Main Methods:
- Benchmarked six backbone architectures (ResNet, EfficientNet, DenseNet, ConvNeXt, ViT) with three fusion strategies.
- Utilized the Chinese Mammography Database (CMMD) with five-fold patient-level cross-validation.
- Reported both side-level and patient-level metrics, including statistical analyses like DeLong's paired AUC test.
Main Results:
- Side-level AUC was consistently higher than patient-level AUC by an average of 17.5 percentage points.
- Vision Transformer (ViT) underperformed Convolutional Neural Networks (CNNs) despite having more parameters.
- Patient-level BI-RADS evaluation with standard aggregation yielded a degenerate macro-AUC of 0.000.
- High prevalence in the cohort made identifying non-malignant cases challenging at the patient level.
Conclusions:
- Evaluation granularity and reporting methodology are critical confounds in mammography classification research.
- Absolute performance metrics from specific datasets should not be extrapolated to population screening settings.
- Future studies must report both side-level and patient-level metrics with consistent rules and robust statistical validation.

