Related Experiment Video
Updated: Apr 19, 2026

07:15
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
7.7K
ONCO-RADS-guided Large Language Models for Extraction and Classification of Incidental Findings on Whole-Body Imaging
Mickael Tordjman1, Murat Yuce1, Zelong Liu1
1BioMedical Engineering and Imaging Institute, Icahn School of Medicine at Mount Sinai, 1470 Madison Ave, New York, NY 10029.
Radiology. Imaging Cancer
|April 17, 2026
Summary
Large language models (LLMs) enhanced by ONCO-RADS significantly improved incidental finding classification in whole-body imaging reports. This strategy outperformed zero-shot LLMs and medical NER models.
Area of Science:
- Radiology and Medical Imaging
- Artificial Intelligence in Healthcare
- Natural Language Processing
Background:
- Incidental findings in whole-body (WB) imaging require accurate extraction and classification for optimal patient management.
- Large Language Models (LLMs) show promise in analyzing unstructured radiology reports, but their performance on incidental findings needs evaluation.
- The Oncologically Relevant Findings Reporting and Data System (ONCO-RADS) provides a standardized framework for reporting critical findings.
Purpose of the Study:
- To assess the performance of LLM-based strategies for extracting and classifying incidental findings from WB imaging reports.
- To specifically evaluate LLM strategies that incorporate the ONCO-RADS framework for improved accuracy and standardization.
- To compare the efficacy of different LLM approaches, including fine-tuned models, zero-shot LLMs, and reference-guided prompting.
Main Methods:
- Retrospective analysis of WB MRI reports from two centers (internal and external datasets).
- Evaluation of ONCO-RADS classification reproducibility among radiologists.
- Comparison of three LLM strategies: fine-tuned medical NER, zero-shot LLMs (ChatGPT-o1, Gemini-2.5-Pro), and ONCO-RADS-guided LLMs.
Main Results:
- ONCO-RADS classification showed excellent interobserver reproducibility (Cohen κ = 0.87).
- ONCO-RADS-guided LLMs achieved significantly higher per-report accuracies (95.6% for ChatGPT-o1, 86.7% for Gemini-2.5-Pro) compared to zero-shot LLMs and medical NER.
- These findings were consistent across both internal and external datasets, including various WB imaging modalities.
Conclusions:
- Reference-guided prompting of LLMs (ChatGPT-o1, Gemini-2.5-Pro) using ONCO-RADS substantially enhances the extraction and classification of incidental findings in WB imaging reports.
- ONCO-RADS-guided LLMs demonstrate superior performance over standard medical NER and zero-shot LLM approaches.
- This integrated strategy offers a robust method for improving the analysis of incidental findings in large-scale imaging data.

