Related Experiment Video For electronic health record
Updated: Mar 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Prompting large language models and evaluating inter- and intra-rater agreement for cancer progression assessment
T Ó Kristjánsson1, A F Henriksen1, M L Hansen1
1Rigshospitalet, Copenhagen, Denmark.
Background:
Manual annotation of free-text radiology reports is time-consuming and costly, delaying real-world evidence (RWE) studies in oncology. This study aimed to evaluate the performance of large language models (LLMs) in annotating cancer progression from Danish free-text radiology reports. The objectives were to determine whether human-to-LLM inter-rater agreement was non-inferior to human-to-human agreement, establish human intra-rater agreement, and develop a framework for tuning LLM performance to RWE needs.
Materials And Methods:
We identified 376 radiology reports from 184 patients with metastatic breast cancer from Danish electronic health records. Six human annotators, including two experts, classified radiology reports as progressive disease (PD) or non-PD. A 'reverse questioning' strategy was used to evaluate five LLM model series (Mistral, Gemma, Gemma 2, Llama 3, and Llama 3.1). Bootstrapping estimated confidence intervals (CIs) and assessed non-inferiority of the best-performing LLM ensemble compared with human agreement, using a non-inferiority margin of 0.1.
Results:
The LLM framework was non-inferior to human annotators with a mean Cohen's kappa of 0.82 (95% CI 0.74-0.89) for human-to-LLM versus 0.79 (95% CI 0.71-0.86) for human-to-human agreement (P < 0.001). The best-performing ensemble model, Llama 3.1:70B, achieved 100% sensitivity, a specificity of 90%, and an F1 score of 84% on the test set. The mean human intra-rater variability was 0.87.
Conclusions:
The proposed LLM framework was non-inferior to human annotators in classifying cancer progression from free-text radiology reports. This offers significant potential for using LLMs as a tool for identifying tumor progression events in clinical assessment and research.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Tumor Progression
Colon cancer is one of the best-documented examples of tumor progression. Early mutation in the APC gene in colon cells causes a small growth on the colon wall called a polyp. With time, this polyp grows into a benign, pre-cancerous tumor. Further...
Cancer Survival Analysis