Related Experiment Video
Updated: May 14, 2026

06:04
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
Can ChatGPT Replace Human Clinical Coders? A Comparative Study in Otology Billing.
Armo Derbarsegian1, Adam S Vesole1, Daniel Q Sun1
1Department of Otolaryngology-Head and Neck Surgery, University of Cincinnati College of Medicine, USA.
The Annals of Otology, Rhinology, and Laryngology
|May 13, 2026
Summary
Large language models like ChatGPT show potential for automating Current Procedural Terminology (CPT) code generation in otologic surgery, but require further refinement for complex cases and accurate billing.
Area of Science:
- Medical informatics
- Artificial intelligence in healthcare
- Surgical coding and billing
Background:
- Accurate Current Procedural Terminology (CPT) code assignment is crucial for medical billing and reimbursement in otologic surgery.
- Manual coding is time-consuming and prone to errors, necessitating exploration of automated solutions.
Purpose of the Study:
- To evaluate the utility of large language models (LLMs), specifically ChatGPT (versions 3.5 and 4), in analyzing operative notes for CPT code generation.
- To compare the performance of ChatGPT in assigning CPT codes against human clinical coders in an otology practice.
Main Methods:
- 191 otologic operative notes from a single surgeon were analyzed using ChatGPT-3.5 and ChatGPT-4.
- Performance was assessed by comparing generated CPT codes to existing billing data, calculating exact and partial match rates, sensitivity, specificity, and work Relative Value Unit (wRVU) differences.
Main Results:
- ChatGPT-4 achieved 14% exact and 33% partial CPT code matches, while ChatGPT-3.5 achieved 22% exact and 32% partial matches.
- High sensitivity and specificity were observed for cochlear implantation (CI) coding (ChatGPT-4: 96% sensitivity, 92% specificity).
- Performance was significantly lower for complex procedures like cartilage grafting, and both models tended to underbill wRVUs compared to human coders.
Conclusions:
- ChatGPT demonstrates potential for automating CPT code assignment in otologic surgery, particularly for procedures like cochlear implantation.
- Current LLM performance is limited by challenges with complex cases, modifier application, and accurate wRVU valuation.
- Further development and validation are needed to optimize LLMs for reliable and comprehensive medical billing automation.