Related Experiment Video
Updated: Jan 8, 2026

A Bioluminescent and Fluorescent Orthotopic Syngeneic Murine Model of Androgen-dependent and Castration-resistant Prostate Cancer
Published on: March 6, 2018
Large language models for toxicity extraction in oncology trials: A real-world benchmark in prostate radiotherapy
Federico Mastroleo1, Mariana Borras-Osorio2, Shiv Patel2
1Department of Radiation Oncology, Mayo Clinic, Rochester, MN, United States; Division of Radiation Oncology, IEO, European Institute of Oncology, IRCCS, Milan, Italy; Department of Oncology and Hemato-Oncology, University of Milan, Milan, Italy.
Background:
Accurate toxicity assessment is critical in oncology trials, yet current reporting frameworks such as the Common Terminology Criteria for Adverse Events (CTCAE) remain labor-intensive and subject to inter-observer variability. Large language models (LLMs) offer potential to automate extraction and grading of adverse events from clinical notes and patient-reported outcomes (PROs), but their comparative performance and cost-effectiveness remain underexplored.
Methods:
We evaluated five off-the-shelf LLMs (Gemini 2.0 Flash, Gemini 2.5 Flash, Gemini 2.5 Pro, GPT-4o, and GPT-5) using a rule-augmented few-shot prompting strategy to extract CTCAE-graded gastrointestinal and genitourinary toxicities from a prospective prostate radiotherapy trial (NCT02874014; n = 55 patients, 8968 toxicity records). Binary and grade-level accuracy, precision, recall, specificity, F1 score, Cohen's kappa, and computational costs were assessed.
Results:
All models achieved high binary accuracy (84.6-87.4 %) and moderate grade accuracy (79.1-82.3 %). GPT-4o reached the best binary (87.4 %) and grade (83.5 %) accuracy, while Gemini 2.5 Pro demonstrated highest sensitivity (74.0 %). Specificity peaked with GPT-4o (96.0 %). Cohen's kappa values indicated moderate agreement (0.552-0.560 for binary; 0.401-0.465 for grades). Costs for the entire extraction varied substantially: Gemini 2.0 Flash delivered competitive accuracy at $0.77 total, whereas Gemini 2.5 Pro and GPT-5 exceeded $21.
Conclusions:
Off-the-shelf LLMs can extract clinically relevant toxicities with performance approaching human inter-rater reliability, at variable but often negligible costs. While grade-level accuracy remains limited, LLM integration into oncology workflows is feasible, offering scalable, low-cost support for toxicity monitoring and data abstraction in clinical research.
More Related Videos
05:39Author Spotlight: Radiotherapy and Clonogenic Assays for Advancing Cancer Research and Personalized Medicine
Published on: April 5, 2024
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018