Related Experiment Video
Updated: Jun 14, 2025

10:42
A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
6.5K
Evaluating the Efficacy of Large Language Models in CPT Coding for Craniofacial Surgery: A Comparative Analysis
Emily L Isch1, Advith Sarikonda2, Abhijeet Sambangi2
1Department of General Surgery, Thomas Jefferson University.
The Journal of Craniofacial Surgery
|September 2, 2024
Summary
Large Language Models show promise for surgical Current Procedural Terminology (CPT) coding. While accuracy varies, AI chatbots offer a resource-efficient alternative to manual coding, especially when primed with operative notes.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Surgery
Background:
- Large Language Models (LLMs) are advancing surgical disciplines.
- Increased interest exists in using LLMs for Current Procedural Terminology (CPT) coding in surgery.
- Manual CPT coding is complex, time-consuming, and faces coder scarcity, necessitating innovative solutions.
Purpose of the Study:
- To evaluate the effectiveness of publicly available LLMs in accurately identifying CPT codes for craniofacial procedures.
- To compare the performance of different AI models in CPT code identification.
Main Methods:
- An observational study assessed 5 LLMs: Perplexity.AI, Bard, BingAI, ChatGPT 3.5, and ChatGPT 4.0.
- A consistent query format was used, including detailed procedure components.
- Responses were classified as correct, partially correct, or incorrect against established CPT codes.
Main Results:
- No significant overall association between AI model type and CPT code correctness was found.
- ChatGPT 4.0 demonstrated higher accuracy for complex CPT codes.
- Perplexity.AI and Bard showed greater consistency with simple CPT codes.
Conclusions:
- AI chatbots offer a promising, accessible, and resource-efficient method for CPT coding in craniofacial surgery, reducing administrative burden.
- While accuracy may be lower than specialized algorithms, LLMs present an attractive alternative due to ease of use.
- Priming AI models with operative notes may enhance accuracy, suggesting a practical strategy for improving CPT coding workflows.

