Related Experiment Video
Updated: Oct 14, 2025

05:58
Evaluation of a Point-of-Care Testing Analyzer for Measuring Peripheral Blood Leukocytes
Published on: March 22, 2022
4.2K
External validation of Machine Learning models for COVID-19 detection based on Complete Blood Count
Andrea Campagner1, Anna Carobene2, Federico Cabitza1
1DISCo, Università degli Studi di Milano-Bicocca, Milan, Italy.
Health Information Science and Systems
|November 1, 2021
Summary
Machine learning models using Complete Blood Count (CBC) data can effectively identify COVID-19 patients. These models offer a rapid, cost-effective alternative to RT-PCR testing, demonstrating strong diagnostic performance and cross-site transportability.
Area of Science:
- Medical Diagnostics
- Computational Biology
- Infectious Disease Research
Background:
- Real-time reverse transcription polymerase chain reaction (rRT-PCR) for COVID-19 diagnosis faces challenges including long turnaround times, reagent shortages, high false-negative rates, and significant costs.
- Routine hematochemical tests, such as Complete Blood Count (CBC), present a faster and more economical alternative for disease diagnosis.
- Machine Learning (ML) approaches are being explored to leverage hematological parameters for developing rapid diagnostic tools, yet external validation of these models remains limited.
Purpose of the Study:
- To externally validate six advanced ML diagnostic models trained on CBC data for COVID-19 detection.
- To assess the real-world applicability and performance of ML models in diverse clinical settings.
- To compare the diagnostic accuracy and calibration of ML models against established methods.
Main Methods:
- External validation of six state-of-the-art ML diagnostic models utilizing Complete Blood Count (CBC) parameters.
- Models were initially trained on a dataset of 816 COVID-19 positive cases.
- Validation was conducted using two independent datasets from different hospitals in northern Italy, comprising 163 and 104 COVID-19 positive cases, evaluating error rates and model calibration.
Main Results:
- The validated ML models achieved an average Area Under the Curve (AUC) of 95% and an average Brier score of 0.11, surpassing existing ML methods.
- The top-performing model, Support Vector Machine (SVM), reported an average AUC of 97.5% (Sensitivity: 87.5%, Specificity: 94%), demonstrating performance comparable to RT-PCR.
- The SVM model also exhibited excellent calibration and good cross-site transportability across different hospital datasets.
Conclusions:
- Externally validated ML models based on CBC data show significant potential for the early identification of COVID-19 patients.
- These models offer a rapid and cost-effective diagnostic solution, complementing traditional testing methods.
- The demonstrated cross-site transportability suggests the broad clinical utility of these ML-based diagnostic tools in various healthcare settings.

