Related Experiment Video
Updated: Jul 27, 2025

Author Spotlight: Bridging Gaps in Anatomy and Establishing a Foundation for Algorithmic Studies
Published on: December 15, 2023
Systematic Analysis of Common Factors Impacting Deep Learning Model Generalizability in Liver Segmentation
Brandon Konkel1, Jacob Macdonald1, Kyle Lafata1
1From the Department of Radiology (B.K., J.M., K.L., I.H.Z., E.B., M.C., G.J., W.F.W., M.R.B.), Department of Radiation Oncology (K.L.), and Department of Medicine, Division of Gastroenterology (M.R.B.), Duke University School of Medicine, Duke University Medical Center, Box 3808, Durham, NC 27710; Department of Electrical & Computer Engineering, Duke University Pratt School of Engineering, Durham, NC (K.L., Y.W.); Department of Radiology, Faculty of Medicine, Benha University, Benha, Egypt (I.H.Z.); Department of Radiology, College of Medicine-Tucson, University of Arizona, Tucson, AZ (E.B.); and Department of Radiology, Rutgers Health-Newark Beth Israel Medical Center, Newark, NJ (M.C.).
Diversifying training data improves deep learning liver segmentation generalizability. Models trained on varied soft-tissue contrast data, like dynamic MRI and opposed-phase scans, perform better across different vendors, MRI types, and CT modalities.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computer-Aided Diagnosis
Background:
- Deep learning models for liver segmentation require robust generalizability across diverse imaging data.
- Variations in imaging protocols, vendors, and modalities (MRI, CT) pose challenges to model performance.
- Understanding the impact of training data characteristics is crucial for developing reliable segmentation tools.
Purpose of the Study:
- To evaluate how different types of training data influence the generalizability of deep learning models for liver segmentation.
- To identify specific imaging sequences that enhance model performance across unseen data domains.
Main Methods:
- A retrospective study utilized 860 abdominal MRI and CT scans, plus 210 public dataset volumes.
- Five single-source models were trained on distinct MRI sequence types (dynportal, dynpre, opposed, ssfse, t1nfs).
- A multisource model (DeepAll) was trained on a diverse dataset combining all sequence types; all models were tested on 18 unseen target domains.
Main Results:
- Single-source models showed varied generalizability; dynamic T1-weighted models performed well on similar data (DSC=0.848).
- The opposed-phase model generalized moderately to unseen MRI types (DSC=0.703) and CT (DSC=0.744).
- The DeepAll model demonstrated strong generalizability across vendors, modalities, and MRI types, outperforming single-source models.
Conclusions:
- Domain shift in liver segmentation is linked to soft-tissue contrast variations.
- Diversifying training data with varied soft-tissue representations effectively bridges domain gaps.
- Multisource training strategies enhance the robustness and generalizability of deep learning liver segmentation models.

