Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models Improve Operative Note Coding Accuracy and Financial Outcomes in Neurotology
Stephanie Younan1, Pearl Doan1, Vanessa S Reyes1
1Department of Otolaryngology-Head and Neck Surgery, University of California-San Francisco, San Francisco, California, USA.
Objective:
To compare coding accuracy and financial impact between an institution's large language model (LLM) and centralized human coders for neurotology operative notes.
Methods:
This retrospective cohort study reviewed 124 consecutive operative notes performed between July 1, 2024, and June 30, 2025, from neurotology attendings at a tertiary academic medical center. Each note was independently coded by the institution's LLM and the centralized coding team. A surgeon-adjudicated reference standard was established through blinded surgeon review and validated by a second blinded neurotologist (Cohen's κ = 0.88). Primary outcomes were coding accuracy and financial variance measured in relative value units (RVUs).
Results:
The LLM achieved significantly higher coding accuracy than human coders (86.3% vs. 49.2%; p = 1.04 × 10-11, McNemar exact test). Human coders demonstrated a mean negative RVU variance of -5.04 (SD 10.79), indicating systematic under-coding, compared with a mean LLM variance of +0.93 (SD 4.85; p < 0.01). Relative human under-coding averaged 17.3% of reference standard RVUs and scaled with procedural complexity. All 61 human errors involved missing or incorrect codes, whereas LLM errors were split between missing and extraneous codes. Extrapolated to the 5-surgeon division, human under-coding projected an annual loss of 1950 RVUs (-$145,342).
Conclusion:
The LLM demonstrated significantly higher concordance with the surgeon-adjudicated reference standard than centralized human coders by 37.1 percentage points. Human errors were driven by under-coding of complex procedures, resulting in substantial projected revenue loss. A hybrid model using LLM-generated drafts verified by specialty-trained coders may optimize coding accuracy and revenue integrity for subspecialty surgical practices.
Level Of Evidence:
N/A.
