Related Experiment Video
Updated: Aug 29, 2026

Automated Midline Shift and Intracranial Pressure Estimation based on Brain CT Images
Published on: April 13, 2013
Systematic evaluation of foundation models for organ-level classification on CT scans
Jan Tagscherer1, Sarah de Boer2, Fennie van der Graaf2
1Department of Medical Imaging, Radboudumc, Nijmegen, The Netherlands. jan.tagscherer@radboudumc.nl.
Purpose:
Foundation models are increasingly used as frozen feature extractors for CT classification tasks, yet the determinants of their downstream performance remain unclear. We assess how foundation model choice, feature aggregation strategy, and abnormality type are associated with performance in organ-level abnormality classification. We further evaluate whether more expressive aggregation strategies outperform simple pooling, and whether performance varies by abnormality type.
Methods:
We evaluate five state-of-the-art medical imaging foundation models on binary organ-level abnormality classification (normal vs. abnormal) across six abdominal organs. Local patch-level embeddings are aggregated using six strategies, including simple pooling methods (e.g., mean pooling) and attention-based multiple instance learning. A linear classifier is trained on feature embeddings, and performance is evaluated on an annotated test set of 200 CT scans.
Results:
Among the evaluated models, 3D CT-native models generally outperformed 2D multi-modal models. None of the tested aggregation strategies significantly outperformed mean pooling, including attention-based multiple instance learning (best-performing aggregation vs. mean: AUC = 0.008, 95% CI ). In our organ-level classification setting, performance differed between abnormality type for all well-performing models (SPECTRE, TAP-CT, CT-FM), with lower AUCs observed for focal abnormalities compared to diffuse abnormalities (largest difference: AUC = 0.108, 95% CI [0.080, 0.135]). This gap was not reduced by the evaluated aggregation strategies.
Conclusion:
In our experiments, downstream performance varied more across foundation models than across aggregation strategies. The choice of aggregation strategy, including attention-based multiple instance learning, did not significantly impact performance in this setting. The persistent gap for focal abnormalities suggests that current representations may insufficiently encode localized disease patterns, motivating the development of localization-aware pre-training approaches.
More Related Videos
Related Concept Videos
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Imaging Studies for Cardiovascular System V: CT
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
Imaging Studies III: Computed Tomography
