Related Experiment Video
Updated: Nov 7, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Impact of diagnosis code grouping method on clinical prediction model performance: A multi-site retrospective
Aman Kansal1, Michael Gao2, Suresh Balu1
1Duke University School of Medicine, Durham, NC, USA; Duke Institute for Health Innovation, Durham, NC, USA.
Objective:
The primary purpose of this work is to systematically assess the performance trade-offs on clinical prediction tasks of four diagnosis code groupings: AHRQ-Elixhauser, Single-level CCS, truncated ICD-9-CM codes, and raw ICD-9-CM codes.
Materials And Methods:
We used two distinct datasets from different geographic regions and patient populations and train models for three prediction tasks: 1-year mortality following an ICU stay, 30-day mortality following surgery, and 30-day complication following surgery. We run multiple commonly-used binary classification models including penalized logistic regression, random forest, and gradient boosted trees. Model performance is evaluated using the Area Under the Receiver Operating Characteristic (AUROC) and the Area Under the Precision-Recall Curve (AUCPR).
Results:
Single-level CCS, truncated codes, and raw codes significantly outperformed AHRQ-Elixhauser ICD grouping when predicting 30-day postoperative complication and one-year mortality after ICU admission. The performance across groupings was more similar in the 30-day postoperative mortality prediction task.
Discussion:
Single-level CCS groupings represent aggregations of raw codes into meaningful clinical concepts and consistently balance interoperability between ICD-9-CM and ICD-10-CM while maintaining strong model performance as measured by AUROC and AUCPR. Key limitations include experimentation across two datasets and three prediction tasks, which although were well labeled and sufficiently prevalent, do not encompass all modeling tasks and outcomes.
Conclusion:
Single-level CCS groupings may serve as a good baseline for future models that incorporate diagnosis codes as features in clinical prediction tasks. Code and a compute environment summary are provided along with the analyses to enable reproducibility and to support future research.
More Related Videos
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Bias in Epidemiological Studies
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Confounding in Epidemiological Studies

