Related Experiment Videos
American College of Surgeons NSQIP Hospital Benchmarking Using Bayesian Variational Inference to Adjust for Many CPT
Yaoming Liu1, Mark E Cohen1, Arielle Grieco1
1. Division of Research and Optimal Patient Care, American College of Surgeons, Chicago, IL.
Journal of the American College of Surgeons
|August 7, 2026
Summary
Bayesian ADVI offers a faster and more accurate method for risk adjustment in surgical outcomes compared to CatBoost. This approach improves benchmarking by utilizing multiple Current Procedural Terminology codes effectively.
Area of Science:
- * Surgical outcomes research and health services analytics.
- * Application of advanced statistical modeling in healthcare.
Background:
- * Current American College of Surgeons National Surgical Quality Improvement Program (ACS NSQIP) benchmarking uses single Current Procedural Terminology (CPT) codes for risk adjustment.
- * Traditional methods struggle with numerous, sparse CPT codes, while CatBoost (CATB) is computationally intensive and can yield unstable results for rare procedures.
- * Bayesian automatic differentiation variational inference (ADVI) is explored as a more efficient alternative.
Purpose of the Study:
- * To evaluate ADVI as a superior method for CPT-based risk adjustment in ACS NSQIP data.
- * To compare ADVI against CATB in terms of calibration, discrimination, computational efficiency, and impact on hospital benchmarking.
Main Methods:
- * Applied CATB and ADVI to ACS NSQIP data (2019-2023) incorporating multiple CPT codes (up to 21 per patient).
- * Assessed performance using Hosmer-Lemeshow statistic (calibration) and area under the receiver operating characteristic curve (discrimination).
- * Measured computational time and analyzed effects on hospital-level morbidity benchmarking.
Main Results:
- * Both ADVI and CATB showed similar discrimination across 31 outcomes.
- * ADVI demonstrated superior calibration (lower Hosmer-Lemeshow values) and significantly reduced computational time compared to CATB.
- * For large models, ADVI's mean computation time was 9.8 minutes versus CATB's 1447.7 minutes, with ADVI achieving a mean Hosmer-Lemeshow of 125.23 vs. CATB's 764.72.
Conclusions:
- * ADVI is a computationally efficient and statistically robust method for multi-CPT code risk adjustment in ACS NSQIP.
- * ADVI enhances calibration and provides more stable estimates for complex procedure combinations.
- * This improved reliability and scalability can significantly benefit procedure-based surgical benchmarking.