Related Experiment Video
Updated: Sep 26, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.0K
Performance of Convolutional Neural Networks for Polyp Localization on Public Colonoscopy Image Datasets
Alba Nogueira-Rodríguez1,2, Miguel Reboiro-Jato1,2, Daniel Glez-Peña1,2
1CINBIO, Department of Computer Science, ESEI-Escuela Superior de Ingeniería Informática, Universidade de Vigo, 32004 Ourense, Spain.
Diagnostics (Basel, Switzerland)
|April 23, 2022
Summary
Deep learning models for polyp detection show performance decay on new datasets. This highlights challenges in inter-dataset testing for artificial intelligence in colonoscopy.
Area of Science:
- Medical imaging
- Artificial intelligence
- Gastroenterology
Background:
- Colorectal cancer is a common malignancy, with colonoscopy being the standard for polyp detection.
- Artificial intelligence (AI), particularly deep learning, is increasingly used for real-time polyp detection and localization in colonoscopy (CADe systems).
- Machine learning model performance is sensitive to dataset variations, especially in inter-dataset testing scenarios.
Purpose of the Study:
- To evaluate the performance of a previously published AI polyp detection model on ten public colonoscopy image datasets.
- To analyze the model's results in the context of 20 other state-of-the-art publications using the same datasets.
- To assess the impact of inter-dataset testing on AI model generalizability for polyp detection.
Main Methods:
- Testing a deep learning polyp detection model on ten diverse public colonoscopy image datasets.
- Comparing the model's performance (F1-score) in intra-dataset (private) and inter-dataset (public) settings.
- Analyzing and contextualizing results against 20 other published studies using the same public datasets.
Main Results:
- The AI model's F1-score decreased by an average of 13.65% when tested on public datasets compared to its performance on a private test set (0.88 intra-dataset).
- Published research shows an average intra-dataset F1-score of 0.91 for state-of-the-art models.
- These models also experienced performance decay in inter-dataset testing, with an average F1-score of 0.83.
Conclusions:
- AI polyp detection models exhibit significant performance degradation when applied to datasets different from their training data.
- Inter-dataset testing is crucial for evaluating the true generalizability and robustness of AI models in colonoscopy.
- Further research is needed to improve the cross-dataset performance of AI systems for reliable polyp detection in clinical practice.

