Related Experiment Videos
Data Extraction From Oncology Imaging Reports by Large Language Models: A Comparative Accuracy Study
Lea P Passweg1, Johannes M Schwenke2, Christof M Schönenberger2
1Division of Medical Oncology, University and University Hospital Basel, Basel, Switzerland.
JCO Clinical Cancer Informatics
|August 10, 2026
Summary
Locally hosted large language models (LLMs) matched human accuracy for identifying cancer metastasis in German radiology reports. However, LLMs showed lower accuracy than humans in assessing treatment response.
Area of Science:
- Artificial Intelligence in Radiology
- Clinical Natural Language Processing
- Medical Data Extraction
Background:
- Manual data extraction from clinical texts is time-consuming and costly.
- Locally hosted large language models (LLMs) offer a potential privacy-preserving alternative for data extraction.
- The performance of LLMs on non-English clinical data, specifically German radiology reports, is not well-established.
Purpose of the Study:
- To evaluate the accuracy of locally hosted LLMs in determining metastasis status from German radiology reports.
- To assess the accuracy of LLMs in classifying treatment response based on German radiology reports.
- To compare the performance of LLMs against human accuracy in these tasks.
Main Methods:
- A retrospective comparative accuracy study involving five locally hosted LLMs (llama3.3:70b, mistral-small:24b, qwq:32b, qwen3:32b, gpt-oss:120b) and human extractors.
- Ground truth was established by duplicate human extraction and senior oncologist adjudication.
- 400 German radiology reports (CT, MRI, PET) from adult cancer patients were analyzed; 300 used for testing after prompt optimization.
Main Results:
- For metastasis status, all LLMs demonstrated non-inferiority to human accuracy (human: 98.4%). The best performing LLM, gpt-oss:120b, achieved 97.6% accuracy.
- For treatment response assessment, human accuracy was 86.0%. All LLMs were inferior to humans, with gpt-oss:120b achieving 78.3% accuracy.
- The study included 400 reports from 317 patients, with a non-inferiority margin of 5 percentage points.
Conclusions:
- Locally hosted LLMs achieve non-inferior accuracy compared to human performance for classifying metastasis status in German radiology reports.
- LLMs currently exhibit inferior accuracy to humans when assessing treatment response from German radiology reports.
- Further research is needed to improve LLM performance in evaluating treatment response in non-English clinical data.