Related Experiment Video
Updated: Jan 11, 2026

06:18
Author Spotlight: Segmentation and VR for Advanced Neurovascular Interventions
Published on: April 5, 2024
1.5K
Evaluating Artificial Intelligence-Assisted Current Procedural Terminology Coding in Vascular Surgery: A Comparison
Brandon Madris1, Keivan Ranjbar1, Andre Critsinelis1
1Division of Vascular Surgery, Cardiovascular Center, Tufts Medical Center, Boston, MA.
Annals of Vascular Surgery
|November 14, 2025
Summary
Artificial intelligence (AI) models like ChatGPT and Perplexity Pro show potential for Current Procedural Terminology (CPT) coding in vascular surgery, but brief notes improve accuracy, necessitating human oversight for optimal results.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Documentation
Background:
- Evaluating AI performance in medical coding is crucial for optimizing healthcare administration.
- Current Procedural Terminology (CPT) code assignment accuracy impacts billing and reimbursement.
- Assessing AI tools like ChatGPT Plus and Perplexity Pro for CPT coding in specialized fields like vascular surgery is an emerging area of research.
Purpose of the Study:
- To compare the accuracy of ChatGPT Plus and Perplexity Pro in assigning CPT codes for vascular surgery cases.
- To evaluate the impact of different documentation formats (full operative notes vs. brief summaries) on AI CPT coding performance.
- To benchmark AI-generated CPT codes against the established coding practices of a hospital's finance department.
Main Methods:
- Analyzed 120 vascular surgery cases from April 2024, using both full operative notes and brief summaries.
- Tested ChatGPT Plus and Perplexity Pro for CPT code generation from the provided case documentation.
- Assessed AI performance using CPT-level (individual code) and case-level (overall case) accuracy metrics, comparing against finance department codes.
- Employed Cohen's kappa analysis to measure inter-rater agreement between AI models and the reference standard.
Main Results:
- Both AI models performed better with brief operative summaries than full notes.
- ChatGPT showed a 43% improvement in CPT-level match rate and a 77% increase in case-level exact matches with brief notes.
- Perplexity Pro demonstrated improved CPT-level accuracy but a slight decrease in case-level exact matches with brief notes.
- Overall agreement remained fair to moderate, indicating a need for human oversight.
Conclusions:
- Concise, structured documentation significantly enhances AI performance in CPT coding.
- ChatGPT demonstrated greater improvement with structured input compared to Perplexity Pro.
- Human oversight is essential for accurate AI-assisted CPT coding, especially with untrained models.
- Untrained AI models offer substantial opportunities for accuracy improvement with specific training data.

