Related Experiment Video
Updated: Jul 17, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Artificial Intelligence for Automated Current Procedural Terminology Coding in Plastic and Reconstructive Surgery
Daniel Oh1, Jeffrey Cripe2, James Baez-Silva2
1From the Department of Statistics and Data Science, University of California, Los Angeles, CA.
Background:
Current Procedural Terminology (CPT) coding in plastic and reconstructive surgery is complex, time-intensive, and prone to error, particularly for procedures involving multiple CPT codes. Recent studies suggest that out-of-the-box large language models (LLMs) demonstrate inconsistent performance for surgical CPT coding.
Methods:
A retrospective comparative study evaluated CPT code set accuracy across baseline LLMs (OpenAI, Gemini, Anthropic), external professional auditors, and a fine-tuned hybrid LLM system (ProCode), using expert consensus coding as the reference standard. A total of 120 operative reports were analyzed by CPT code set complexity into low (1-2 codes), medium (3 codes), and high (≥4 codes) complexity, with 40 cases per group. Accuracy was defined as exact code set agreement with the reference standard.
Results:
ProCode demonstrated the highest accuracy across all complexity levels (low, 100%; medium, 80%; and high, 80%), outperforming human auditors and baseline LLMs. Baseline LLM accuracy declined noticeably with increasing complexity. Differences in correctness by method were significant overall and within each complexity tier (all P < 0.05).
Conclusions:
A fine-tuned, domain-specific LLM substantially outperforms baseline LLMs and external professional auditors for automated CPT coding in plastic and reconstructive surgery, particularly for complex multicoded procedures. These findings support the role of specialized AI systems in improving coding accuracy and surgical billing workflows.