Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated review of spine surgery operative reports with large language models: a pilot study of GPT reasoning models
Rushmin Khazanchi1, Alan G Soetikno1, Avani Chopra1
1Feinberg School of Medicine, Northwestern University, 420 E Superior St, Chicago, IL 60611, USA.
Abstract:
Clinical documentation burden increased exponentially since the Health Information Technology for Economic and Clinical Health Act in 2009. Large language models (LLMs) offer substantial potential for alleviating this burden by analyzing clinical notes. This study validates GPT-based LLMs for extracting clinical data from spine surgery operative reports, expanding prior work to a new note type and clinical domain. Operative reports from 88 patients undergoing lumbar decompression and/or fusion surgery in 2022 for lumbar spondylolisthesis were retrospectively reviewed. Three independent reviewers for a variety of surgical characteristics (type of procedure, graft usage, hardware usage, etc). Five GPT models (4o, o1, o3 low reasoning, o3medium reasoning, o3 high reasoning) were then prompted to extract the same set of variables. Model performance metrics were calculated and compared to the human ground truth. Reviewer agreement for the presence of certain variables varied considerably (52 %-99 %). No significant performance differences were observed among the models (p > 0.05), although o3 -low had the lowest duration (20.032 sec, p = 0.0070) and cost ($0.008, p < 0.0001). Its accuracy ranged from 75 % to 100 % for reviewed variables. Human agreement was positively correlated with model accuracy (r = 0.67, p = 0.0003) and uncertainty from the model was highly sensitive for errors (ranging from 0.63 to 1.00). Our analysis demonstrates the potential of GPT models in automating manual chart review with a high degree of accuracy, possibly surpassing human reviewers. Further work is necessary to extend and validate these models in other institutions or procedures.
