Related Experiment Video
Updated: Mar 31, 2026

Utilizing 18F-FDG PET/CT Imaging and Quantitative Histology to Measure Dynamic Changes in the Glucose Metabolism in Mouse Models of Lung Cancer
Published on: July 21, 2018
Conceptual proposal for LLM-generated FDG PET/CT follow-up reports in melanoma: a pilot study on model stability and
Wolfram A Bosbach1, Marie S Heide1, Nasir Gözlügöl1
1Department of Nuclear Medicine, Inselspital, Bern University Hospital, University of Bern, Bern, Switzerland.
Purpose:
Oncological patients regularly undergo PET/CT re-staging, which requires a report that outlines their current disease status and highlights relevant changes compared to the previous PET/CT. Large language models (LLMs) may be helpful with documentation in the future. This study is a pilot on LLM performance, focusing on test-retest stability and reproducibility.
Methods:
Three textbook melanoma follow-up cases of increasing complexity (involving one to eight organs) were selected. From standardized text-only prompts (no imaging data), follow-up reports were written by GPT-4o, Claude Sonnet 4 (each producing three independent revisions), and three nuclear medicine residents. This yielded nine reports per case (27 in total). Six blinded nuclear medicine experts (three internal, three external) performed test-retest evaluations of report quality and authorship identification.
Results:
The cosine similarity analysis revealed high intra-case coherence (mean: 0.599-0.727) regardless of authorship. The external human readers consistently rated reports higher than the internal human readers. The LLM-generated reports received comparable or superior ratings to human reports, with Claude achieving the highest external reader scores (mean 0.926, standard deviation 0.263, on a 0-1 scale). Human performance declined with case complexity, while Claude, in particular, improved. The external readers significantly preferred the LLM impressions (Fisher's exact test, p = 0.005). Neither the human nor LLM readers reliably identified authorship (balanced accuracy 0.343-0.500).
Conclusion:
In this pilot, blinded expert evaluation demonstrated that current LLMs can generate reports for melanoma [18F]fluorodeoxyglucose PET/CT of comparable quality to human-authored reports from text prompts in this study. High test-retest stability was obtained. Larger future studies will be required to confirm these findings.
More Related Videos
10:28Gene Regulation and Targeted Therapy in Gastric Cancer Peritoneal Metastasis: Radiological Findings from Dual Energy CT and PET/CT
Published on: January 22, 2018
10:04Analysis of 18FDG PET/CT Imaging as a Tool for Studying Mycobacterium tuberculosis Infection and Treatment in Non-human Primates
Published on: September 5, 2017