Related Experiment Video
Updated: May 31, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
484
Evaluating Large Language Models for Automated CPT Code Prediction in Endovascular Neurosurgery.
Joanna M Roy1, D Mitchell Self1, Emily Isch2
1Department of Neurological Surgery, Thomas Jefferson University, Philadelphia, PA, USA.
Journal of Medical Systems
|January 24, 2025
Summary
Large language models (LLMs) can identify Current Procedural Terminology (CPT) codes from neurosurgery reports. AtlasGPT and ChatGPT showed higher accuracy than Gemini, suggesting potential for healthcare cost reduction.
Area of Science:
- Artificial Intelligence
- Neurosurgery
- Medical Informatics
Background:
- Large language models (LLMs) are increasingly used for healthcare automation, including clinical documentation in neurosurgery.
- Accurate coding of procedures is crucial for billing and healthcare management.
Purpose of the Study:
- To evaluate the ability of three LLMs (ChatGPT 4.0, AtlasGPT, and Gemini) to identify Current Procedural Terminology (CPT) codes from endovascular neurosurgery operative reports.
- To compare the performance of these LLMs in CPT code identification.
Main Methods:
- Thirty endovascular neurosurgery operative reports were analyzed.
- Three LLMs were tasked with identifying CPT codes for diagnostic or interventional procedures.
- Responses were categorized as correct, partially correct, or incorrect.
- Statistical tests (Chi-Square, Kruskal Wallis) were used for comparison.
Main Results:
- AtlasGPT achieved 98.3% partially correct CPT code identification, followed by ChatGPT (86.7%) and Gemini (30%).
- For fully correct CPT codes, AtlasGPT averaged 35.3%, ChatGPT 35.1%, and Gemini 8.9%.
- AtlasGPT and ChatGPT significantly outperformed Gemini in CPT code identification accuracy.
Conclusions:
- Untrained LLMs demonstrate a capability for identifying partially correct CPT codes in endovascular neurosurgery.
- Further training of LLMs could improve CPT code accuracy and potentially reduce healthcare expenditures.

