Related Experiment Video
Updated: Aug 6, 2025

Using Computer-based Image Analysis to Improve Quantification of Lung Metastasis in the 4T1 Breast Cancer Model
Published on: October 2, 2020
VAI-B: a multicenter platform for the external validation of artificial intelligence algorithms in breast imaging
Fernando Cossío1,2, Haiko Schurz1, Mathias Engström3
1Karolinska Institute, Department of Oncology-Pathology, Stockholm, Sweden.
This article introduces VAI-B, a secure, hybrid platform designed to independently test artificial intelligence tools used for breast cancer screening. By combining cloud-based processing with local data storage, the system allows hospitals and developers to evaluate diagnostic software fairly using large, diverse patient datasets.
Area of Science:
- Medical imaging informatics within VAI-B validation research
- Oncology diagnostics and screening technology
Background:
No prior work had resolved the challenge of independently verifying commercial diagnostic software performance across diverse clinical settings. Many vendors currently provide automated tools for cancer detection, yet standardized, transparent testing remains largely absent. That uncertainty drove the creation of robust, external validation frameworks to ensure reliable clinical outcomes. Prior research has shown that proprietary algorithms often lack rigorous assessment on independent datasets before deployment. This gap motivated the development of specialized infrastructure capable of handling massive mammography archives securely. Existing methods frequently rely on internal vendor testing, which may introduce bias or limit generalizability. Researchers recognize that external evaluation is a prerequisite for safe, widespread integration of automated systems into routine practice. The current landscape necessitates a scalable solution that balances computational flexibility with stringent patient privacy requirements.
Purpose Of The Study:
The primary aim of this work is to introduce a specialized platform for the independent assessment of diagnostic software in breast imaging. This initiative addresses the urgent requirement for transparent, external testing of commercial systems used in clinical screening. The authors seek to provide a solution that enables fair benchmarking of various algorithms against standardized, high-quality datasets. They intend to bridge the gap between vendor-led internal testing and the need for objective, real-world performance data. By creating this framework, the team hopes to facilitate safer technology adoption within hospital networks. The researchers also aim to offer a scalable environment that accommodates future growth in both vendor participation and regional data contributions. They focus on balancing the need for massive computational power with the strict privacy constraints of medical records. This study serves as a foundational step toward establishing a reliable, multicenter ecosystem for continuous algorithm evaluation.
Main Methods:
The team designed a hybrid infrastructure that integrates cloud-based services with local, secure data management systems. Their review approach involved establishing a centralized repository at the Karolinska Institute to house sensitive patient information. They utilized a MongoDB database alongside specialized python scripts to organize incoming clinical records. The researchers defined a large case-control population using national quality registries to ensure statistical power. They extracted images and expert assessments from three distinct regional healthcare providers across Sweden. These records were processed by three different commercial systems within a virtual private cloud environment. The investigators calculated abnormality scores to measure the performance of each diagnostic tool against known cancer outcomes. This systematic workflow allowed for the independent evaluation of thousands of examinations without compromising data security.
Main Results:
The platform successfully processed 105,706 mammography examinations during the initial pilot phase. These records originated from a cohort of 8,080 patients with confirmed cancer and 36,339 healthy control subjects. The researchers integrated data from three different regions to ensure a diverse and representative testing environment. Three distinct commercial vendors participated in the evaluation, with their systems generating abnormality scores for every image. The study confirmed that the hybrid architecture could scale effectively to handle high-volume inference tasks. By comparing automated outputs against radiologist assessments, the team established a baseline for future performance monitoring. The database currently maintains a comprehensive record of all processed images and associated clinical outcomes. This pilot demonstrates that the infrastructure is fully operational and capable of supporting large-scale, multicenter validation efforts.
Conclusions:
The authors propose that their hybrid architecture provides a sustainable model for independent software assessment in clinical environments. This synthesis suggests that separating inference processing from data storage effectively addresses privacy concerns while maintaining scalability. The team indicates that their framework facilitates faster development cycles for commercial partners by providing standardized performance feedback. They claim that hospitals benefit from safer adoption practices when using this independent validation approach. The researchers note that the platform remains ready for expansion to include additional vendors or regional healthcare providers. Their findings imply that such infrastructure is necessary for the objective benchmarking of diagnostic tools. The study demonstrates that large-scale, multicenter data integration is feasible for ongoing algorithm monitoring. Finally, the authors conclude that this initiative supports the broader goal of improving breast cancer screening accuracy through transparent, external verification.
Frequently Asked Questions
The platform utilizes a hybrid architecture, combining cloud-based computing for inference tasks with on-premises storage for sensitive patient records. This dual-layered approach allows for scalable processing power while ensuring that raw radiological data remains protected within secure local environments.
The system relies on a MongoDB database and a custom python package to organize and track information. These tools manage the four required inputs: radiological images, AI-generated inferences, expert radiologist assessments, and final cancer outcomes.
The researchers emphasize that on-premises storage is necessary to preserve patient privacy. By keeping raw images in a secure local environment at the Karolinska Institute, the platform avoids exposing sensitive clinical information to the public cloud during the validation process.
The platform requires four distinct data types: radiological images, AI-generated inferences, radiologist assessments, and cancer outcomes. These components collectively enable a comprehensive comparison between automated software predictions and established clinical truths.
The pilot study processed 105,706 mammography examinations from 8,080 cancer patients and 36,339 healthy individuals. This large, multicenter dataset allows for a robust assessment of abnormality scores generated by three different commercial vendors.
The authors propose that this platform enables faster development cycles for vendors and safer adoption for hospitals. By providing an independent, transparent testing environment, the system helps stakeholders identify performance gaps before deploying tools in clinical practice.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018