Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models Improve Operative Note Coding Accuracy and Financial Outcomes in Neurotology
Stephanie Younan1, Pearl Doan1, Vanessa S Reyes1
1Department of Otolaryngology-Head and Neck Surgery, University of California-San Francisco, San Francisco, California, USA.
The Laryngoscope
|July 30, 2026
Summary
A large language model (LLM) significantly improved coding accuracy for neurotology operative notes compared to human coders. This AI-driven approach can prevent substantial revenue loss from under-coding complex procedures.
Area of Science:
- Medical informatics
- Artificial intelligence in healthcare
- Surgical coding and billing
Background:
- Accurate coding of operative notes is crucial for revenue integrity in healthcare.
- Traditional human coding processes can be prone to errors and financial underestimation, particularly in complex surgical subspecialties.
- The integration of artificial intelligence, specifically large language models (LLMs), presents a potential solution to enhance coding efficiency and accuracy.
Purpose of the Study:
- To compare the coding accuracy and financial impact of an LLM against centralized human coders for neurotology operative notes.
- To evaluate the performance of an LLM in identifying and coding complex procedures accurately.
- To quantify the financial implications of coding discrepancies between LLM and human coders.
Main Methods:
- A retrospective cohort study of 124 neurotology operative notes was conducted.
- Notes were independently coded by an institutional LLM and a centralized human coding team.
- A surgeon-adjudicated reference standard, validated by a second neurotologist (Cohen's κ = 0.88), was used to assess primary outcomes: coding accuracy and financial variance (RVUs).
Main Results:
- The LLM achieved significantly higher coding accuracy (86.3%) than human coders (49.2%; p < 1.04 × 10-11).
- Human coders exhibited significant under-coding, with a mean RVU variance of -5.04, projecting an annual loss of 1950 RVUs (-$145,342) for a 5-surgeon division.
- LLM errors were less frequent and involved both missing and extraneous codes, while human errors primarily involved missing or incorrect codes, especially in complex procedures.
Conclusions:
- The LLM demonstrated superior coding accuracy and concordance with the reference standard compared to human coders.
- Human under-coding of complex procedures led to substantial projected revenue loss.
- A hybrid model, utilizing LLM-generated drafts verified by specialty-trained coders, may optimize coding accuracy and revenue integrity in surgical practices.
