Related Experiment Video
Updated: May 5, 2026

14:08
Automated Midline Shift and Intracranial Pressure Estimation based on Brain CT Images
Published on: April 13, 2013
42.8K
Comparative analysis of supervised and self-supervised learning with small and imbalanced medical imaging datasets
Andrea Espis1, Chiara Marzi2, Stefano Diciotti3,4
1Department of Electrical, Electronic, and Information Engineering "Guglielmo Marconi" - DEI, University of Bologna, Via dell'Università 50, 47521, Cesena, Italy.
Scientific Reports
|September 2, 2025
Summary
Supervised learning generally outperforms self-supervised learning on small, imbalanced medical imaging datasets. Careful selection of learning methods is crucial, considering data size and label availability for optimal performance.
Area of Science:
- Computer Vision
- Medical Imaging Analysis
- Machine Learning
Background:
- Self-supervised learning (SSL) shows promise for reducing labeled data needs in computer vision.
- Real-world medical applications often involve limited dataset sizes and imbalanced class distributions, unlike large, balanced datasets (e.g., ImageNet).
Purpose of the Study:
- To compare the performance of SSL and supervised learning (SL) on small, imbalanced medical imaging datasets.
- To evaluate the impact of varying label availability and class frequency distributions on model performance.
Main Methods:
- Experiments were conducted on four binary classification tasks: age prediction and Alzheimer's disease diagnosis (brain MRI), pneumonia detection (chest X-ray), and retinal disease classification (OCT).
- Training set sizes varied, with a mean of 843 (age prediction), 771 (Alzheimer's), 1,214 (pneumonia), and 33,484 (retinal disease) images.
- Model training was repeated with different random seeds to ensure result reliability and assess uncertainty.
Main Results:
- Supervised learning (SL) demonstrated superior performance compared to the evaluated SSL methods across most experiments with small training sets.
- This advantage of SL persisted even when only a limited amount of labeled data was accessible.
- The findings were consistent across diverse medical imaging modalities and classification tasks.
Conclusions:
- The study underscores that supervised learning is often more effective than self-supervised learning for small, imbalanced medical imaging datasets.
- Optimal selection of machine learning paradigms is critical and depends heavily on specific application constraints, including training set size, data balance, and label availability.

