Related Experiment Video
Updated: Aug 23, 2026

Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
A large language models-assisted and expert-corrected workflow for preoperative anesthesia assessment drafts: A
Shuhan Gu1, Jianfan Ping1, Huidan Lin2
1Department of Anesthesiology, The First Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, Zhejiang, China.
Background:
Large language models (LLMs) may help organize clinical information, but their use in perioperative settings requires careful evaluation because errors may have immediate safety implications. This study aimed to describe the feasibility and perceived usefulness of a single-centre, expert-corrected LLM workflow for preparing preoperative anesthesia assessment drafts for complex consultation cases. Secondary aims were to describe error patterns identified by anesthesiologists and to explore residents' perceptions after reviewing expert-corrected materials.
Methods:
We retrospectively selected 15 complex preoperative anesthesia consultation cases from The First Affiliated Hospital, Zhejiang University School of Medicine. Complex cases were defined as cases referred for specialized preoperative anesthesia consultation because of multiple comorbidities, clinically relevant organ dysfunction, implanted devices, major cardiopulmonary or vascular disease, or other features requiring individualized anesthetic assessment. A large language model-based artificial intelligence system, DeepSeek-R1, was used to generate structured draft assessments from de-identified case information and a standardized prompt. The original LLM-generated drafts were independently reviewed by three experienced anesthesiologists for information completeness, scientific plausibility, and focus on key risk factors. All resident-facing documents were then corrected by senior anesthesiologists before being distributed to first-year anesthesia residents, who completed a questionnaire on perceived usefulness, guidance value, educational support, and error recognition. No comparator group, blinding, or objective performance outcome was included.
Results:
The LLM-generated drafts were structurally complete but contained clinically relevant errors, mainly in risk assessment and anesthetic plan formulation. Thirty-two error instances were identified across the 15 cases, including errors related to proposed plans, American Society of Anesthesiologists (ASA) physical status classification, and abnormal test interpretation. After expert correction, residents reported favorable perceptions of the materials, but their self-reported ability to identify expert-annotated errors remained limited. These findings should be interpreted as perception-based observations rather than evidence of educational effectiveness.
Conclusions:
In this single-centre exploratory study, an LLM-assisted workflow was feasible for producing preliminary preoperative anesthesia assessment drafts, but expert correction was essential before resident-facing use. The findings do not establish standalone clinical or educational effectiveness of the LLM. Instead, they highlight the gap between structural completeness and clinical reliability, and suggest that any use of LLM-generated anesthesia assessment drafts should remain adjunctive, supervised, and explicitly framed to reduce automation bias.