Related Experiment Videos
Application of large language models for assigning clear cell likelihood score v2.0 from free-text MRI reports: a
Yanting Duan1, Yilei Zhao2, Maowei He3
1Department of Radiology, Beilun Branch of the First Affiliated Hospital, College of Medicine, Zhejiang University, Ningbo, Zhejiang, China.
Rationale And Objectives:
To evaluate the performance of four large language models (LLMs) for automated feature extraction and clear cell likelihood score version 2.0 (ccLS v2.0) assignment from free-text reports using a multi-step prompting strategy.
Materials And Methods:
This retrospective multicenter study included 1,048 magnetic resonance imaging (MRI) reports of renal masses from three institutions (2020-2026). Four LLMs (DeepSeek-V3.2, Qwen 3.5, GPT-5.3, and Gemini 3.0) were evaluated for automated ccLS v2.0 assignment through a three-stage prompting workflow: (1) feature extraction, (2) rule-based criteria matching, and (3) final categorization. Performance was evaluated against expert radiologist consensus-based ccLS v2.0 assignments. Accordingly, these results primarily reflect agreement in task execution rather than diagnostic accuracy for ccRCC.
Results:
Gemini 3.0 achieved the highest overall accuracy, correctly assigning ccLS categories in 683 of 802 reports in the internal cohort (85.2%; 95% CI, 82.5%-87.5%) and 211 of 246 reports in the external validation cohort (85.8%; 95% CI, 80.9%-89.6%). All LLMs achieved accuracy greater than 80% for assigning high-risk categories (ccLS 4-5). Gemini 3.0 showed comparatively strong performance in extracting individual imaging features, such as T2-weighted imaging (T2WI) hyperintensity. However, all models showed reduced performance for features requiring multi-step interpretation, particularly the arterial-to-delayed enhancement ratio (ADR), with Qwen 3.5 showing the weakest performance for these complex features. In the pathology subgroup of the internal cohort (n = 237), expert report-based scores yielded an AUC of 0.799 for identifying ccRCC, with LLM-assigned scores showing AUCs of 0.768-0.791 and small AUC differences from the expert standard (-0.031 to -0.008) with overlapping 95% CIs.
Conclusion:
LLMs employing multi-step prompting, particularly Gemini 3.0, demonstrated favorable performance in assigning ccLS v2.0 categories from free-text reports. Despite limitations in complex feature reasoning, these models showed potential for supporting automated ccLS assessment in radiology workflows; however, in their current implementation, they should be considered decision-support tools that require expert verification rather than autonomous, unsupervised scoring.