Related Experiment Video
Updated: Feb 28, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
3.6K
Quality versus quantity of training datasets for artificial intelligence-based whole liver segmentation
Austin Castelo1, Caleb O'Connor1, Aashish C Gupta1
1Department of Imaging Physics, The University of Texas MD Anderson Cancer Center, Houston, TX 77030, USA.
Medrxiv : the Preprint Server for Health Sciences
|February 27, 2026
Summary
High-quality, smaller datasets for AI model training in medical imaging can match the performance of much larger, mixed-quality datasets. Dataset curation impacts AI segmentation generalizability, showing nuanced tradeoffs between quality and quantity.
Area of Science:
- Medical imaging and artificial intelligence
- Computational pathology and radiology
- Machine learning for medical image analysis
Background:
- Artificial intelligence (AI) segmentation models require large, curated datasets for effective training, which are often limited in medical applications.
- The quality and quantity of annotated data significantly influence the performance and generalizability of AI models in medical image segmentation.
Purpose of the Study:
- To compare the impact of dataset annotation quality versus quantity on whole liver AI segmentation performance.
- To evaluate how different dataset curation strategies affect the Dice Similarity Coefficient (DSC), surface DSC (SD), Hausdorff distance (HD), and slice DSC (Slice DSC).
Main Methods:
- Trained 3D nnU-Net segmentation models using datasets with varying curation levels (mixed vs. highly curated) and sizes.
- Utilized 3,089 abdominal CT scans, withholding 249 for testing and external validation.
- Evaluated model performance using DSC, SD 2mm, HD95, and Slice DSC metrics.
Main Results:
- A highly curated 244-scan dataset achieved performance statistically indistinguishable from a 2,840-scan mixed-curation dataset on 3D metrics.
- The 710-scan mixed-curation dataset significantly outperformed the highly curated 244-scan model on external validation scans (Slice DSC: 0.929 vs. 0.923).
- Highly curated datasets demonstrated performance equivalent to datasets an order of magnitude larger, with larger mixed-curation datasets showing benefits in generalizability.
Conclusions:
- Tradeoffs between dataset quality and quantity in AI model training are nuanced and depend on specific goals.
- High-quality annotations can yield comparable performance to larger, less curated datasets for certain segmentation tasks.
- Larger, mixed-curation datasets may offer advantages in model generalizability and local performance improvements.

