Related Experiment Video
Updated: Jun 30, 2026

Observational Study Protocol for Repeated Clinical Examination and Critical Care Ultrasonography Within the Simple Intensive Care Studies
Published on: January 16, 2019
Racial/Ethnic Disparities in Severity Scoring Systems for Sepsis Patients in Intensive Care Units in the USA:
Peili Liu1,2, Xiaoqi Li1, Yanyan Zhao3
1School of Public Health, Cheeloo College of Medicine, Shandong University, Jinan, 250012, China.
Objectives:
This study aims to evaluate the performance disparities of commonly used severity scoring systems among sepsis patients of different ethnicities.
Methods:
We analyzed data from 51,519 sepsis patients in three routine hospital datasets from the United States. We evaluated the discrimination and calibration of three severity scoring systems across four ethnic groups (White, Black, Hispanic, and Asian). These three scoring systems were the Oxford Acute Severity of Illness Score (OASIS), Logistic Organ Dysfunction System (LODS), and Simplified Acute Physiology Score II (SAPS II). The primary outcome was in-hospital mortality. Discrimination was assessed using the area under the receiver operating characteristic (AUROC) curve, and calibration was analyzed using the standardized mortality ratio (SMR).
Results:
The discrimination of the three scoring systems was generally stable across different ethnic groups. For instance, the AUROC of SAPS II ranged from 0.76 to 0.77 in Medical Information Mart for Intensive Care III (MIMIC-III), 0.74 to 0.78 in Medical Information Mart for Intensive Care IV (MIMIC-IV), and 0.71 to 0.72 in eICU Collaborative Research Database (eICU-CRD). However, significant biases were observed in calibration performance. Specifically, the three scoring systems generally overestimated mortality (SMR < 1) in most settings, although calibration patterns varied across databases. This overestimation was particularly pronounced in Black or Hispanic patients, with statistical significance in some cases when compared to White or Asian patients (adjusted p-value < 0.05).
Conclusions:
Severity scoring systems generally tended to overestimate mortality across different ethnic groups, although the direction and magnitude of calibration bias varied across settings. This suggests that caution is needed when using scoring systems for clinical decision-making, especially for patients from different ethnic backgrounds.

