Related Experiment Video
Updated: Sep 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy of Retrieval-Augmented Large Language Model-Generated Preconsult Summaries in Breast Surgical Oncology
Tracy-Ann Moo1, Robert James1,2, Solange Bayard1
1Department of Surgery, Breast Service, Memorial Sloan Kettering Cancer Center, New York, NY.
Purpose:
Radiology and pathology reports are important in breast surgical oncology planning, but are often unstructured and variable, requiring time-intensive previsit review and preparation. Here we evaluate the accuracy, completeness, and safety of retrieval-augmented large language model (LLM)-generated structured preconsult summaries compared with clinical staff-authored summaries.
Methods:
In this single-institution retrospective validation study, 200 randomly sampled new breast surgical oncology consultations were included. LLM summaries were generated and compared with human summaries documented in consultation notes. Field-level accuracy, omission rate, and fabrication rate among extracted radiology and pathology variables were measured. Error analysis included patient level and clinically significant fabrication rates.
Results:
Artificial intelligence (AI)-generated summaries demonstrated high field-level accuracy, with fabrication rates ≤2%. Accuracy rates were significantly higher than those observed in staff-authored notes across multiple domains, including nodal status (95% v 72%, P < .001) and documentation of invasive tumor component on biopsy (94% v 30%, P < .001). In contrast, staff-authored summaries were more accurate for receptor status (96% v 86%, P = .002). At the patient level, fabrication occurred in 16 AI-generated summaries (8%), including six clinically significant cases (3%). Staff-authored summaries demonstrated fabrication in six cases (3.0%), with one clinically significant instance (0.5%).
Conclusion:
In this retrospective validation study, structured retrieval-augmented LLM-generated preconsult summaries demonstrated field-level accuracy comparable with or exceeding staff-authored documentation across most radiologic and pathologic variables, with low per-field fabrication rates. Clinically significant errors occurred in a small proportion of summaries and were largely confined to identifiable high-risk variables that are amenable to targeted verification. With appropriate human oversight and safeguards, LLM-based structured information extraction may support documentation standardization and improved workflow efficiency in breast surgical oncology.