Related Experiment Video
Updated: Aug 6, 2025

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Effect of Dataset Size and Medical Image Modality on Convolutional Neural Network Model Performance for Automated
Harrison C Gottlich1, Adriana V Gregory2, Vidit Sharma3
1Mayo Clinic Alix School of Medicine, Mayo Clinic, Rochester, MN, USA.
Determining optimal medical image dataset size is crucial for AI performance. This study introduces an exponential-plateau model to predict the number of training images needed for maximum segmentation accuracy across different modalities and structures.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Machine Learning
Background:
- Accurate medical image segmentation is vital for diagnosing and treating diseases like renal tumors.
- The optimal size of training datasets for deep learning models in medical imaging remains a challenge.
- Previous studies have not established a clear method for predicting dataset size requirements for peak model performance.
Purpose of the Study:
- To investigate the efficacy of an exponential-plateau model for determining the optimal training dataset size for medical image segmentation.
- To establish the relationship between dataset size and model generalizability performance across different imaging modalities and target structures.
- To provide a predictive tool for researchers to ascertain when additional training data will not significantly improve model performance.
Main Methods:
- Retrospective collection of CT and MR images of patients with renal tumors.
- Assembly of modality-based datasets with varying image counts (50-300) for model training and validation (80-20 split).
- Evaluation against a held-out test set (50 images) and analysis using an exponential-plateau model to identify performance plateaus.
Main Results:
- The exponential-plateau model successfully predicted performance plateaus for segmenting non-neoplastic kidney regions and tumor regions on CT and MR images.
- Specific numbers of training-validation images required to reach performance plateaus were identified for different segmentation tasks (e.g., 54 for non-neoplastic kidney on CT, 389 for tumor on MR).
- Experiments with the KiTS21 dataset and different model architectures (nn-UNet 2D/3D) confirmed that modality, target structure, and architecture influence required dataset size.
Conclusions:
- The developed exponential-plateau modeling approach effectively predicts the dataset size needed to achieve maximum performance in medical image segmentation.
- The required number of training images varies significantly based on imaging modality, the specific anatomical structure being segmented, and the chosen model architecture.
- This methodology offers a valuable tool for researchers to optimize data acquisition and model training, preventing unnecessary data collection and improving efficiency.
Related Concept Videos
Imaging Studies I: Kidney, Ureter, and Bladder Studies
Imaging Studies for Cardiovascular System V: CT
Imaging Studies IV: Magnetic Resonance Imaging
Imaging Studies III: Computed Tomography
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...

