Deep learning-based PSMA PET segmentation repeatability: A post-hoc analysis of a single-center, prospective,
Jake Kendrick1,2,3, Roslyn J Francis4,5,6,7, Ghulam Mubashar Hassan8,4
1School of Physics, Mathematics and Computing, The University of Western Australia, 35 Stirling Highway, Mailbag M013, Crawley, Perth, WA, 6009, Australia. jake.kendrick@uwa.edu.au.
Objectives:
The primary aim of this study was to quantify the subject-level test-retest repeatability of artificial intelligence (AI)-derived PSMA PET imaging biomarkers using a previously developed model. The secondary aim was to assess the performance of this segmentation model, which was trained on [68 Ga]Ga-PSMA-11 PET scans, on [18F]F-PSMA-1007 PET scans.
Methods:
This was a post-hoc analysis of a prospective, single-center, test-retest trial. Seventeen patients with metastatic prostate cancer (mPCa) were randomised into groups, either receiving the same tracer ([68 Ga]Ga-PSMA-11 or [18F]F-PSMA-1007) for both scans (intra-tracer group, n = 9) or a different tracer (inter-tracer group, n = 8). Scans were delineated using a fully automated AI method and semi-automatically. The subject-level repeatability of four imaging biomarkers, including PSMA-positive tumour volume, was quantified.
Results:
Repeatability analysis demonstrated poorer repeatability for all biomarkers in the inter-tracer group. In the intra-tracer group, the AI-derived PSMA-positive tumour volume had a repeatability coefficient of 13.8% for higher volume disease patients (≥ median tumour volume). There was no significant difference in the per-scan lesion-level positive predictive value of the AI model between [68 Ga]Ga-PSMA-11 and [18F]F-PSMA-1007 PET scans (0.88, IQR 0.69-1.00 vs. 0.78, IQR 0.54-1.00, p = 0.60).
Conclusion:
AI-based PSMA-positive tumour volume calculations have repeatability limits that are consistent with the use of the Response Evaluation Criteria in PSMA PET/CT (RECIP 1.0) criteria for higher volume disease patients when the same tracer is used. Substantially wider repeatability limits in the inter-tracer group provide evidence that response assessment should be conducted using the same radiotracer.


