Related Experiment Video
Updated: Jun 23, 2025

An Orthotopic Resectional Mouse Model of Pancreatic Cancer
Published on: September 24, 2020
Large Language Models for Automated Synoptic Reports and Resectability Categorization in Pancreatic Cancer
Rajesh Bhayana1, Bipin Nanda1, Taher Dehkharghanian1
1From University Medical Imaging Toronto, Joint Department of Medical Imaging, University Health Network, Princess Margaret Cancer Centre, Department of Medical Imaging, University of Toronto, Toronto General Hospital, 200 Elizabeth St, Peter Munk Building, 1st Fl, Toronto, ON, Canada M5G 24C (R.B., B.N., T.D., S.K.); Department of Biostatistics (Y.D.) and HPB Surgical Oncology (C.G.S., C.A.M., D.H., S.G.), University Health Network, Toronto, Ontario, Canada; Departments of Medicine (N.B., G.E., D.D.) and Surgery (C.G.S., C.A.M., D.H., S.G.), University of Toronto, Toronto, Ontario, Canada; and Department of Radiology, Massachusetts General Hospital, Harvard Medical School, Boston, Mass (A.K.).
Large language models (LLMs) like GPT-4 can create accurate synoptic radiology reports for pancreatic cancer, improving surgical decision-making. AI-generated reports enhance surgeon accuracy and efficiency in assessing tumor resectability.
Area of Science:
- Artificial Intelligence in Radiology
- Oncology
- Medical Informatics
Background:
- Structured radiology reports for pancreatic ductal adenocarcinoma (PDAC) enhance surgical decision-making compared to free-text reports.
- Radiologist adoption of structured reporting is inconsistent, and resectability criteria application varies.
- Large language models (LLMs) offer potential for automating structured report generation.
Purpose of the Study:
- To assess LLM performance in automatically generating PDAC synoptic reports from original radiology reports.
- To evaluate LLM capabilities in categorizing tumor resectability based on CT findings.
- To compare surgeon accuracy and efficiency using AI-generated versus original reports.
Main Methods:
- Retrospective analysis of 180 PDAC CT reports.
- GPT-3.5 and GPT-4 were prompted to extract 14 key findings and categorize resectability using various strategies.
- Radiologist and surgeon review compared AI-generated reports against original reports for accuracy and time.
Main Results:
- GPT-4 significantly outperformed GPT-3.5 in generating synoptic reports (F1 score: 0.997 vs 0.967).
- GPT-4 demonstrated higher precision in extracting critical features like superior mesenteric artery involvement.
- GPT-4 with chain-of-thought prompting achieved 92% accuracy in resectability categorization, surpassing other methods.
- Surgeons were more accurate (83% vs 76%) and efficient (58% time saving) using AI-generated reports.
Conclusions:
- GPT-4 can generate near-perfect PDAC synoptic reports, addressing the need for structured reporting.
- Chain-of-thought prompting enhances GPT-4's accuracy in classifying tumor resectability.
- AI-generated reports improve surgical assessment accuracy and efficiency for PDAC.

