Related Experiment Video
Updated: Jun 6, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Large Language Models for Extraction of International Classification of Diseases Codes from Clinical
Large language models (LLMs) show poor performance in extracting International Classification of Diseases-tenth revision - clinical modification (ICD-10-CM) codes from inpatient notes. Human coders remain superior for accurate ICD-10-CM code extraction in clinical documentation.
Area of Science:
- Health Informatics
- Artificial Intelligence in Healthcare
- Clinical Documentation Improvement
Background:
- Accurate International Classification of Diseases-tenth revision - clinical modification (ICD-10-CM) code extraction is crucial for healthcare reimbursement and coding.
- Automating ICD-10-CM code extraction from clinical documentation has faced significant challenges.
- This study addresses the need for improved automated coding solutions.
Purpose of the Study:
- To evaluate the performance of various large language models (LLMs) in extracting ICD-10-CM codes from unstructured inpatient clinical notes.
- To benchmark the accuracy of LLMs against human medical coders.
- To identify the strengths and weaknesses of current LLMs in this specific task.
Main Methods:
- Compared six LLMs (GPT-3.5, GPT-4, Claude 2.1, Claude 3, Gemini Advanced, Llama 2-70b) against a human coder.
- Utilized deidentified inpatient notes from American Health Information Management Association Vlab authentic patient cases.
- Employed a standardized prompt for LLM code extraction and a 3M Encoder with 2022 ICD-10-CM Coding Guidelines for the human coder.
Main Results:
- Analyzed 50 inpatient notes (23 H&Ps, 27 progress notes).
- Human coder identified a median of 4 ICD-10-CM codes per note.
- LLMs extracted a median of 5-11 codes per note, with GPT-4 showing the best performance but only 15.2% overall agreement with the human coder.
Conclusions:
- Current large language models demonstrate inadequate performance for extracting ICD-10-CM codes from inpatient clinical notes when compared to human coders.
- Further advancements are needed to improve LLM accuracy for automated clinical coding.
- Human expertise remains essential for reliable ICD-10-CM code assignment.
More Related Videos
Related Concept Videos
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Statistical Software for Data Analysis and Clinical Trials
Formulating and Validating Nursing Diagnosis I
There are thirteen domains...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Nursing Interventions II: Selecting and Classifying the Nursing Interventions

