Related Experiment Video
Updated: May 10, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Conversion of Mixed-Language Free-Text CT Reports of Pancreatic Cancer to National Comprehensive Cancer Network
Hokun Kim1, Bohyun Kim2, Moon Hyung Choi3
1Department of Radiology, Seoul St. Mary's Hospital, College of Medicine, The Catholic University of Korea, Seoul, Republic of Korea.
Generative pre-trained transformer-4 (GPT-4) models show high accuracy in creating structured reports from mixed-language pancreatic cancer CT scans. Human radiologist oversight remains crucial for resectability assessments.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing in Radiology
- Oncology Decision Support
Background:
- Pancreatic ductal adenocarcinoma (PDAC) diagnosis relies on detailed CT imaging reports.
- Mixed-language (English/Korean) narrative reports pose challenges for automated analysis.
- Standardized structured reporting (SR) is crucial for consistent clinical decision-making.
Purpose of the Study:
- To assess the feasibility of GPT-4 Turbo and GPT-4o in generating structured reports (SRs) from mixed-language PDAC CT reports.
- To evaluate the accuracy of GPT-4 models in categorizing PDAC tumor resectability based on NCCN guidelines.
- To compare the performance of GPT-4 Turbo and GPT-4o in these tasks.
Main Methods:
- Retrospective analysis of 115 mixed-language pancreas-protocol CT reports (English/Korean) from two institutions (Jan 2021-Dec 2023).
- GPT-4 Turbo and GPT-4o models were prompted to generate SRs and classify tumor resectability.
- Model outputs were compared against a radiologist-derived reference standard, with triplicates and majority voting for final output.
Main Results:
- GPT-4 Turbo and GPT-4o achieved comparable accuracy in SR generation (92.3% and 92.2%, respectively).
- GPT-4 Turbo demonstrated significantly higher accuracy in resectability categorization (81.7%) compared to GPT-4o (67.0%).
- Primary errors for GPT-4 Turbo involved inaccurate data extraction in SR generation and violation of resectability criteria.
Conclusions:
- GPT-4 Turbo and GPT-4o show promise for generating NCCN-based structured reports from mixed-language PDAC CT narratives.
- Human radiologist review is indispensable for accurate resectability determination due to model limitations.
More Related Videos
11:18Generation of Comprehensive Thoracic Oncology Database - Tool for Translational Research
Published on: January 22, 2011
07:47Author Spotlight: Unveiling Transmembrane Protein Family-Related Markers in Gastric Cancer and Implications for Targeted Therapies
Published on: September 15, 2023