Hepatocellular Carcinoma Segmentation on MRI: Five-Fold ATLAS Benchmarking, Phase-Wise External Stress Testing, and
1East Lancashire Hospitals NHS Trust, Royal Blackburn Teaching Hospital, Haslingden Road, Blackburn BB2 3HH, UK.
Background/Objectives:
To establish an auditable nnU-Net v2 HCC MRI segmentation benchmark, quantify phase- and cohort-dependent transportability without external fine-tuning, evaluate false-positive safety on HCC-negative examinations, and test whether complementary arterial-plus-venous input improves performance in OpenSwissHCC.
Methods:
A 3D full-resolution nnU-Net was trained by five-fold cross-validation on 60 public ATLAS cases. The unchanged five-fold ensemble was applied separately to arterial, portal/venous, delayed and native T1-weighted images in LiverHccSeg and OpenSwissHCC. All 132 OpenSwissHCC subjects (63 HCC-positive and 69 HCC-negative) underwent inference in every phase; positive-case overlap and detection were reported separately from negative-examination specificity. A patient-paired five-fold ablation compared venous-only with arterial-plus-venous OpenSwissHCC training. MRI-to-CT application to 97 valid-reference HCC-TACE-Seg cases was exploratory.
Results:
The ATLAS out-of-fold mean Dice was 0.5315 (95% CI 0.4478-0.6126). The OpenSwissHCC positive mean Dice was 0.3388 arterial, 0.1267 venous, 0.1810 delayed and 0.0814 native; the corresponding detection was 55.6%, 28.1%, 27.6% and 23.8%, while the negative-examination specificity ranged from 50.7% to 81.2%. The detection was lower when all of the documented HCCs were <20 mm. Arterial-plus-venous training increased the mean Dice from 0.2269 to 0.3705 (paired difference +0.1435, 95% CI +0.0520 to +0.2343; p = 0.00223) and reduced the false-positive voxel burden (p = 0.0394), although the binary detection increase was not significant by exact McNemar testing (p = 0.0639). The MRI-to-CT mean Dice was 0.0286.
Conclusions:
The HCC segmentation performance was strongly domain-sensitive. Multiphasic input improved segmentation quality within the cohort, but reliable clinical use requires substantially broader development and prospective validation.


