Related Experiment Videos
Deep learning-based segmentation and detection of focal liver lesions on multi-sequence non-contrast MRI
Duoduo Zhang1, Ke Wang1, Pengsheng Wu2
1Department of Radiology, Peking University First Hospital, Beijing, China.
Purpose:
To evaluate the performance of a deep learning-based artificial intelligence (AI) model for the multi-sequence automated segmentation and detection of focal liver lesions (FLLs) on non-contrast MRI (NC-MRI).
Methods:
A total of 1010 NC-MRI were included as the training set, and 144 consecutive MRI were collected as the hold-out set. The publicly available LLDMMRI-2023 and CHAOS datasets were utilized for external validation. T1WI, fat-saturated T2WI (fsT2WI), and high-b-value DWI (DWI_High) were included. Radiologist annotations on individual sequences served as the reference standard. A two-stage sequential 3D V-Net was implemented for FLLs segmentation. Segmentation performance was assessed using Dice similarity coefficient (DSC). A lesion was considered detected if the DSC between model-predicted mask and reference standard exceeded zero. Detection performance was evaluated using sensitivity at the lesion and patient level, and false-positive lesion per sequence. For verified sensitivity, a lesion on one sequence was considered a true positive if detected on any other sequence for the same patient. Subgroup analyses were performed according to lesion size. Lesions were first split using a 20 mm cutoff, then classified into four size groups: <10 mm, 10-20 mm, 20-40 mm, and ≥ 40 mm. Subgroup analyses comparing benign and malignant lesions were performed in the external validation set.
Results:
Mean DSC for FLLs segmentation were 0.535, 0.590 and 0.596 for the training, hold-out and external validation sets. Sensitivities in the hold-out set ranged from 0.794 to 0.882 across T1WI, fsT2WI, and DWI_High, with all verified sensitivities > 0.93. In the external validation set, sensitivities were 0.688 (0.640-0.732), 0.864 (0.833-0.891) and 0.918 (0.889-0.941), while verified sensitivities reached 0.978, 0.978 and 0.972, respectively. Per-sequence false positive lesions were 0.50, 1.94, 1.42 for the hold-out set and 0.60, 1.66, 0.90 for the external validation set. Patient-level sensitivity was 0.988 and 0.987 across the two datasets. In the external validation set, DWI_High showed a statistically significant difference in sensitivity between lesions in < 20 mm and ≥ 40 mm subgroups (p < 0.001). For fsT2WI, benign lesions achieved superior detection sensitivity, with significant differences in 10-20 mm (p < 0.001) and 20-40 mm subgroups (p < 0.05). For DWI_High, malignant lesions yielded higher sensitivity in < 10 mm subgroups (p < 0.001).
Conclusion:
The deep learning-based AI model enables automated segmentation and detection of FLLs on NC-MRI, with acceptable performance across different lesion sizes.