Related Experiment Video
Updated: May 23, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of Large Language Models for Automated Simple CPT Coding in Foot and Ankle Surgery.
Eve R Glenn1, Ariana Rowshan1, Eric Mao1
1Department of Orthopaedic Surgery, Johns Hopkins University, Baltimore, Maryland, USA.
Foot & Ankle Orthopaedics
|May 22, 2026
Summary
Large language models (LLMs) show variable accuracy in generating Current Procedural Terminology (CPT) codes for foot and ankle surgery. While some LLMs show promise as aids, they are not yet reliable for independent clinical use in CPT coding.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Orthopaedic Surgery Coding
Background:
- Accurate Current Procedural Terminology (CPT) coding is essential for billing and reimbursement in foot and ankle surgery.
- Current coding practices are often time-consuming and error-prone.
- Large Language Models (LLMs) present a potential solution to automate coding and reduce administrative burdens in medicine.
Purpose of the Study:
- To evaluate the accuracy of five publicly available LLMs in generating CPT codes for simple foot and ankle procedures.
- To compare the performance of ChatGPT-5 Mini, Google Gemini 2.5 Flash, Claude 4.0 Sonnet, Deepseek V3, and Perplexity.
Main Methods:
- Twenty-one common single-CPT-code foot and ankle procedures were selected.
- Each LLM was queried four times using standardized prompts for CPT code generation.
- Accuracy was determined by the correct identification of CPT codes, with statistical analysis comparing model performance.
Main Results:
- Perplexity demonstrated the highest accuracy at 92.9%, while Deepseek V3 performed the worst at 48.2%.
- Significant differences in coding accuracy were observed among the evaluated LLMs (P < .001).
- Pairwise comparisons indicated Perplexity, Google Gemini 2.5 Flash, and Claude 4.0 Sonnet outperformed Deepseek V3 and/or ChatGPT-5 Mini.
Conclusions:
- LLM performance in CPT coding for simple foot and ankle procedures is highly variable.
- Current LLMs are not sufficiently reliable for independent clinical application in CPT coding.
- Certain LLMs may function as preliminary aids when supplemented by thorough human verification.
