Related Experiment Video
Updated: Feb 12, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Validation of 13 102 International Classification of Diseases, Tenth Revision, Clinical Modification codes using a
Yichen Wang1,2, Yilin Song3, Rex Siu4
1Division of Gastroenterology and Hepatology, Department of Medicine, Mayo Clinic, Jacksonville, FL 32224, United States.
International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) codes show high accuracy, especially for principal diagnoses. A large language model (LLM) system effectively validates these codes, offering a scalable solution for improving data quality.
Area of Science:
- Medical Informatics
- Health Data Science
- Clinical Coding Systems
Background:
- Accurate medical coding is crucial for healthcare administration, billing, and research.
- International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) codes are widely used but require validation.
- Large language models (LLMs) present a novel approach for automated data validation.
Purpose of the Study:
- To evaluate the accuracy of ICD-10-CM codes for diverse diagnoses.
- To assess the performance of a GPT-4o-based large language model (LLM) system in validating ICD-10-CM codes.
- To compare LLM-based validation with traditional methods.
Main Methods:
- Retrospective analysis of hospital admissions from the MIMIC-IV database.
- Development and refinement of a GPT-4o LLM system for ICD-10-CM code validation.
- Calculation of positive predictive values (PPV) for principal and secondary diagnoses, and LLM performance metrics (accuracy, sensitivity, specificity).
Main Results:
- Overall ICD-10-CM code PPV was 84.6%, with principal diagnoses at 93.9% and secondary diagnoses at 83.8%.
- The LLM system achieved 93.6% accuracy, 95.4% sensitivity, and 85.2% specificity in code validation.
- PPV decreased with later diagnosis positions, and many secondary diagnoses represented historical conditions or post-admission changes.
Conclusions:
- ICD-10-CM codes demonstrate high overall accuracy, with notable variations based on diagnosis position and condition type.
- The validated LLM system performs comparably to physician review, offering a scalable solution for enhancing coding accuracy.
- Integrating LLM-based auditing into clinical workflows can significantly improve the quality of administrative and research data.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
06:16Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
Related Concept Videos
lncRNA - Long Non-coding RNAs
Histone Modification
Acetylation
The enzyme histone acetyltransferase adds acetyl group to the histones. Another enzyme, histone...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Chronic Kidney Disease II: Clinical Manifestations
Reliability and Validity
Internal Energy