Related Experiment Videos
High AUROC can mask decision failure in sepsis transcriptomic classifiers: Preprocessing stability outweighs post hoc
1Quanzhou Medical College, Quanzhou, Fujian, People's Republic of China.
Background:
High AUROC is often taken as evidence that a transcriptomic classifier is promising, but rank discrimination can conceal fixed-threshold failure after cohort or platform transfer.
Methods:
We benchmarked four GEO whole-blood cohorts: GSE65682 for discovery, GSE95233 for external microarray validation, GSE154918 for cross-platform RNA-seq validation, and GSE28750 for non-infectious inflammation stress testing. We compared logistic-regression workflows using training-derived standard scaling, training-derived robust scaling, sample-wise rank normalization with training-derived scaling, and robust scaling using unsupervised external-cohort reference statistics. Internal performance used five-fold cross-validation with fold-contained imputation and scaling. External uncertainty used 2,000 stratified bootstrap replicates.
Results:
Internal discrimination was very high for all strategies, but external validation revealed threshold collapse for training-derived standard and robust scaling. In GSE154918, both had balanced accuracy 0.50 at the 0.5 threshold despite very high AUROC, equivalent to random classification at that fixed threshold. The strict-inductive sample-rank strategy preserved fixed-threshold performance across external cohorts (balanced accuracy 0.95-1.00). Robust external-cohort adaptation also performed well (0.95-1.00) but uses unlabeled external-cohort distribution statistics and is therefore reported as adaptation rather than fixed single-sample transfer. Calibration and regularization sensitivity did not rescue the failing training-derived scaling strategies. In the sepsis-versus-non-infectious-inflammation stress test, robust external-cohort adaptation had the highest observed balanced accuracy (0.80, 95% CI 0.61-0.95), but its difference from sample-rank normalization was uncertain in paired bootstrap analysis.
Conclusions:
High AUROC can mask fixed-threshold failure in sepsis transcriptomic classifiers. In this benchmark, strict-inductive sample-rank normalization was the most stable fixed external strategy, while robust external-cohort scaling was best interpreted as unsupervised cohort adaptation whose reliability depends on external reference-sample availability.