Related Experiment Video
Updated: Jan 27, 2026

Pathological Analysis of Lung Metastasis Following Lateral Tail-Vein Injection of Tumor Cells
Published on: May 20, 2020
Evaluating reproducibility of AI algorithms in digital pathology with DAPPER
Andrea Bizzego1,2, Nicole Bussola1,3, Marco Chierici1
1Fondazione Bruno Kessler, Trento, Italy.
We developed DAPPER, a framework for validating AI models in digital pathology. This ensures reproducible results by analyzing variability in predictive biomarkers, crucial for AI in healthcare.
Area of Science:
- Digital pathology and Artificial Intelligence (AI)
- Computational pathology and machine learning
- Biomarker validation and reproducibility
Background:
- Artificial Intelligence (AI) is transforming healthcare, particularly in digital pathology, by enhancing computer vision tasks.
- Deep learning models can extract valuable features from pathology images, potentially improving diagnostic accuracy and identifying novel histological patterns.
- Ensuring the reproducibility and accuracy of AI-driven predictive models in digital pathology is a critical challenge.
Purpose of the Study:
- To introduce and validate the DAPPER (Data Analysis Plan for Predictive biomarker Evaluation and Reproducibility) framework for assessing the accuracy and reproducibility of deep learning models in digital pathology.
- To analyze the sources of variability in predictive biomarkers derived from digital pathology images.
- To provide a standardized approach for evaluating AI models in digital pathology, promoting robust and reliable clinical applications.
Main Methods:
- The DAPPER framework was applied to evaluate predictive models for tissue of origin identification using 787 Whole Slide Images from the Genotype-Tissue Expression (GTEx) project.
- Three deep learning architectures (VGG, ResNet, Inception) were used as feature extractors, coupled with three classifiers (multilayer perceptron, Support Vector Machine, Random Forests).
- Model performance was assessed across datasets with varying numbers of classes (5, 10, 20, 30) and diagnostic tests were employed to detect selection bias and reproducibility risks.
Main Results:
- The study analyzed the accuracy and feature stability of various machine learning classifiers applied to deep learning features from digital pathology images.
- Diagnostic tests, such as random label analysis, were demonstrated to be essential for identifying potential selection bias and risks to reproducibility in AI models.
- The DAPPER framework and associated software, including the HINT benchmark dataset, were released to facilitate standardization and validation in AI for digital pathology.
Conclusions:
- The DAPPER framework provides a rigorous methodology for validating AI predictive models in digital pathology, addressing key challenges in reproducibility.
- Standardized validation is essential for the reliable integration of AI tools into routine pathology reporting and clinical trials.
- The released DAPPER software and HINT dataset serve as valuable resources for the research community to advance AI in digital pathology.
More Related Videos
Related Concept Videos
Trial and Error and Algorithm
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Social Foundations of Self IV: Self in Digital Communication
Self-Evaluation: Self-Enhancement and Self-Verification
Nursing Evaluation
Self-Evaluation Maintenance Model

