Related Experiment Video
Updated: Jul 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Artificial intelligence for procedural coding in cardiac critical care: Evaluating large language models for current
Leon Fan1, Sukethram Sivakumar2, Ernesto Marin1
1Division of Cardiac Surgery, Department of Surgery, Johns Hopkins University School of Medicine, Baltimore, MD, USA.
Large Language Models show potential for automating critical care procedural coding, but struggle with complex cases like extracorporeal membrane oxygenation (ECMO). Further domain-specific fine-tuning is needed for accurate clinical application.
Area of Science:
- Critical Care Medicine
- Health Informatics
- Artificial Intelligence in Healthcare
Background:
- Accurate procedural coding is vital for resource allocation, billing, and quality reporting in critical care.
- Manual coding in high-acuity settings like cardiovascular surgical intensive care units (CVSICU) is complex and error-prone, particularly for procedures such as extracorporeal membrane oxygenation (ECMO).
- Large Language Models (LLMs) present a potential scalable solution for automating procedural coding.
Purpose of the Study:
- To systematically evaluate the performance of six publicly accessible LLMs in assigning Current Procedural Terminology (CPT) codes to procedures performed in a CVSICU.
- To compare the accuracy of different LLMs for both general CVSICU procedures and complex ECMO-related interventions.
Main Methods:
- Six LLMs (GPT-4, Claude 3.7 Sonnet, Perplexity, DeepSeek, Google Gemini 2.5 Pro, Mistral) were tested.
- Models were prompted to assign CPT codes to 47 CVSICU procedures, including 7 ECMO interventions, from a single tertiary center.
- Code accuracy was evaluated, and statistical comparisons were performed for inter-model performance differences.
Main Results:
- For non-ECMO procedures, Gemini 2.5 Pro and Perplexity achieved the highest accuracy (88%).
- For ECMO-related codes, Perplexity demonstrated the highest accuracy (86%), followed by Gemini 2.5 Pro (71%).
- Significant inter-model performance differences were observed, with GPT-4.0 showing 0% accuracy for ECMO codes.
Conclusions:
- LLMs like Perplexity and Gemini show promise for automated procedural coding in critical care.
- A key limitation is the models' current difficulty in understanding context-dependent nuances, especially for ECMO procedures.
- Future research should focus on domain-specific fine-tuning of LLMs to improve accuracy in high-acuity clinical settings.
More Related Videos
07:46Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
Published on: August 9, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025