Related Experiment Video
Updated: Aug 6, 2026

07:15
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Evaluating Locally Deployed Large Language Models for Ga-68 PSMA PET/CT Report Mining in Prostate Cancer
Jagrati Chaudhary1, Param Dev Sharma2, Sanjay Kumar3
1Department of Nuclear Medicine, All India Institute of Medical Sciences, New Delhi, Ansari Nagar, 110029 India.
Nuclear Medicine and Molecular Imaging
|July 26, 2026
Summary
Large Language Models (LLMs) like Llama 3 and Gemma 2 can quickly extract key diagnostic features from Ga-68 PSMA PET/CT reports. While effective for initial screening, their performance on rare findings suggests human-assisted workflows are optimal.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing for Clinical Reports
- Radiomics and Data Extraction
Background:
- Unstructured radiology reports pose challenges for data extraction.
- Large Language Models (LLMs) offer potential for automated data interpretation.
- Ga-68 PSMA PET/CT is crucial for prostate cancer diagnosis and management.
Purpose of the Study:
- To evaluate the zero-shot extraction capabilities of Llama 3.2:3b and Gemma 2:2b models.
- To assess the accuracy of LLMs in identifying key diagnostic features from Ga-68 PSMA PET/CT reports.
- To determine the feasibility of converting free-text reports into structured, queryable data.
Main Methods:
- Retrospective analysis of 50 de-identified Ga-68 PSMA PET/CT reports.
- Expert annotation and inter-reader agreement for ground truth establishment.
- Batch processing with standardized prompts for Llama 3.2:3b and Gemma 2:2b models.
- Benchmarking model outputs against adjudicated dual-expert consensus.
Main Results:
- Both Llama 3.2:3b and Gemma 2:2b demonstrated rapid and reliable feature extraction.
- Excellent inter-reader agreement (average κ = 0.882) was achieved.
- Llama 3.2:3b showed higher sensitivity (83.7%) and NPV (97.3%).
- Gemma 2:2b achieved higher overall accuracy (86.2%) and specificity (88.5%).
Conclusions:
- Open-weight LLMs Llama 3.2:3b and Gemma 2:2b facilitate rapid extraction of PET/CT findings, significantly faster than manual review.
- Llama 3.2:3b excels in sensitivity, while Gemma 2:2b leads in specificity.
- Performance varies with feature prevalence and complexity.
- Models are best suited for initial screening or human-assisted workflows due to limitations in rare findings and positive predictive value.
