Related Experiment Video
Updated: Jul 16, 2026

03:07
Single-Port Robotic-assisted Transaxillary Breast-conserving Surgery: A Prospective, Single-arm, Non-randomized Phase IIa Clinical Trial
Published on: August 19, 2025
Artificial Intelligence-Generated Electronic Medical Record Summarization in Breast Surgical Oncology
Ko Un Park1,2,3,4, Bergen K Sather5,6, Anika Shah7
1Division of Breast Surgery, Department of Surgery, Brigham and Women's Hospital, Boston, MA, USA. kpark16@bwh.harvard.edu.
Annals of Surgical Oncology
|July 15, 2026
Summary
Retrieval-Augmented Generation (RAG)-enabled GPT-4o AI summaries show promise in breast oncology but require human review. While accurate and useful, summaries sometimes lack thoroughness and contain critical errors, necessitating further development.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Oncology Informatics
Background:
- Reviewing external oncology records is time-consuming due to varied formats.
- A Retrieval-Augmented Generation (RAG)-enabled GPT-4o summarization agent was developed to address this challenge.
- The study evaluated the agent's impact on clinical workflows and summary quality in breast surgical oncology.
Purpose of the Study:
- To assess the quality of AI-generated summaries of external oncology records.
- To evaluate the impact of a RAG-enabled GPT-4o agent on clinical workflows.
- To determine the accuracy, usefulness, and potential errors in AI-generated summaries.
Main Methods:
- Initial evaluation of a GPT-4o/RAG agent on 50 oncologic reports.
- Prospective pilot test using a modified Provider Documentation Summarization Quality Instrument (PDSQI-9).
- Analysis included error criticality, user-reported errors, and pre/post-use surveys on documentation burden (NASA TLX).
Main Results:
- AI summaries rated highly for accuracy, usefulness, succinctness, and source citation.
- Thoroughness was rated low in 45% of summaries; 40% contained errors, with 52% deemed critical.
- The most common error type involved imaging (68%); perceived time savings were neutral, but qualitative feedback noted benefits for simple cases.
Conclusions:
- RAG-enabled GPT-4o summaries are favorably rated but frequently lack thoroughness and can contain critical errors.
- Human review remains essential for ensuring summary accuracy and patient safety.
- Further technological iteration is required before widespread implementation in clinical practice.