Related Experiment Video
Updated: Feb 12, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Validation of 13 102 International Classification of Diseases, Tenth Revision, Clinical Modification codes using a
Yichen Wang1,2, Yilin Song3, Rex Siu4
1Division of Gastroenterology and Hepatology, Department of Medicine, Mayo Clinic, Jacksonville, FL 32224, United States.
Objectives:
To comprehensively evaluate the validity of International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) codes for both prevalent diagnoses and less common diseases, and to assess the performance of a large language model (LLM)-based system in validating these codes.
Materials And Methods:
This retrospective study analyzed hospital admissions from the Medical Information Mart for Intensive Care IV (MIMIC-IV) database. We developed a validated LLM-based system using GPT-4o, refined through iterative prompt engineering, to assess ICD-10-CM code validity. We measured the positive predictive value (PPV) of ICD-10-CM codes, PPV of principal and secondary diagnoses, and the performance of an LLM-based system in code validation.
Results:
Among 865 079 assigned codes, the PPV was 84.6% (95% CI, 84.5%-84.6%). Principal diagnoses had a PPV of 93.9% (95% CI, 93.7%-94.1%), while secondary diagnoses had a PPV of 83.8% (95% CI, 83.7%-83.9%). The LLM system demonstrated high performance in validating ICD codes, achieving 93.6% accuracy, 95.4% sensitivity, and 85.2% specificity. Among correctly assigned secondary diagnoses, the majority (67.9%) represented historical or baseline conditions, while 32.1% reflected active conditions that deviated from baseline status; 22.3% of these emerged after hospital admission. PPV decreases with later diagnosis positions, with the largest decline occurring between principal and secondary diagnoses.
Discussion And Conclusion:
In this large-scale evaluation, ICD-10-CM codes exhibited generally high accuracy, though variability existed by position and condition type. A validated LLM system performed comparably to physician review and offers a scalable means to improve coding accuracy. These findings support the potential for integrating LLM-based auditing into routine workflows to strengthen the quality of administrative and research data.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
06:16Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
Related Concept Videos
lncRNA - Long Non-coding RNAs
Histone Modification
Acetylation
The enzyme histone acetyltransferase adds acetyl group to the histones. Another enzyme, histone...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Chronic Kidney Disease II: Clinical Manifestations
Reliability and Validity
Internal Energy