Related Experiment Videos
MoMBS: Mixed-order sampling improves training on heterogeneous-quality data for universal lesion detection
Han Li1, Jingsong Liu2, Peter Schüffler2
1Chair of Computer Aided Medical Procedures, Technical University of Munich, Germany; Medical Imaging, Robotics, Analytic Computing & Learning (MIRACLE) Lab, YRD-RIGHT, USTC Suzhou Institute for Advanced Research, Suzhou, Jiangsu, 215123, China; Institute of Pathology, Technical University of Munich, Germany; Munich Center for Machine Learning (MCML), Munich, Germany.
A new Mixed-order Minibatch Sampling (MoMBS) method improves universal lesion detection by better identifying challenging training images. This approach enhances model performance by prioritizing underrepresented samples, leading to more accurate diagnoses in computed tomography scans.
Area of Science:
- Medical Imaging Analysis
- Machine Learning for Healthcare
- Computer-Aided Diagnosis
Background:
- Universal lesion detection (ULD) in computed tomography (CT) is crucial for computer-aided diagnosis.
- ULD faces challenges due to significant variations in image quality and labeling accuracy.
- Existing methods like self-paced learning (SCL) and ordered hard example mining (OHEM) inadequately capture sample difficulty.
Purpose of the Study:
- To propose a novel minibatch sampling strategy, Mixed-order Minibatch Sampling (MoMBS), for ULD.
- To address the limitations of existing methods in handling heterogeneous training data quality.
- To improve the accuracy and robustness of ULD models.
Main Methods:
- Introduced Mixed-order Minibatch Sampling (MoMBS), a novel approach for minibatch sampling.
- MoMBS utilizes a joint measure of sample loss and uncertainty for finer-grained difficulty categorization.
- Prioritizes underrepresented samples in minibatch gradient contribution, inspired by human learning.
Main Results:
- MoMBS demonstrated significant improvements over two state-of-the-art (SOTA) methods on the DeepLesion dataset for ULD (0.97%-7.28% gains).
- Validated generalizability across diverse tasks: Seg-19 (up to 2.3% improvement), CIFAR100-LT (up to 5.3% improvement), and CIFAR100-NL (up to 4.6% improvement).
- Consistently outperformed existing SOTA approaches in all evaluated settings.
Conclusions:
- MoMBS offers a more accurate estimation of sample difficulty compared to loss-based methods.
- The proposed sampling strategy effectively handles noisy and overfitted examples, improving model training.
- MoMBS represents a significant advancement in training strategies for ULD and other machine learning tasks facing data heterogeneity.