Related Experiment Video
Updated: Aug 14, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Application of a population-based severity scoring system to individual patients results in frequent
Frank V Booth1, Mary Short, Andrew F Shorr
1Eli Lilly and Company, Indianapolis, IN, USA. boothfv@lilly.com
Introduction:
APACHE II (AP2) was developed to allow a systematic examination of intensive care unit outcomes in a risk adjusted manner. AP2 has been widely adopted in clinical trials to assure broad consistency amongst different groups. Although errors in calculating the true AP2 score may not be reducible below 15%, the self-canceling effect of random errors reduces the importance of such errors when applied to large populations. It has been suggested that a threshold AP2 score be used in clinical decision making for individual patients. This study reports the AP2 scoring errors of researchers involved in a large sepsis trial and models the consequences of such an error rate for individual severe sepsis patients.
Methods:
Fifty-six researchers with explicit training in data abstraction and completion of the AP2 score received scenarios consisting of composites of real patient histories. Descriptive statistics were calculated for each scenario. The standard deviations were calculated compared with an adjudicated score. Intraclass correlations for inter-observer reliability were performed using Shrout-Fleiss methodology. Theoretical distribution curves were calculated for a broad range of AP2 scores using standard deviations of 6, 9 and 12. For each curve, the misclassification rate was determined using an AP2 score cut-off of >or=25. The percentage of misclassifications for each true AP2 score was then applied to the corresponding AP2 score obtained from the PROGRESS severe sepsis registry.
Results:
The error rate for the total AP2 score was 86% (individual variables were in the range 10% to 87%). Intraclass correlation for the inter-observer reliability was 0.51. Of the patients from the PROGRESS registry. 50% had AP2 scores in the range 17 to 28. Within this interquartile range, 70% to 85% of all misclassified patients would reside.
Conclusion:
It is more likely that an individual patient will be scored incorrectly than correctly. The data obtained from the scenarios indicated that as the true AP2 score approached an arbitrary cut-off point of 25, the observed misclassification rate increased. Integrating our study of AP2 score errors with the published literature leads us to conclude that the AP2 is an inappropriate sole tool for resource allocation decisions for individual patients.
Insights
The APACHE II (AP2) scoring system frequently misclassifies individual severe sepsis patients, especially near the 25-point threshold. This high error rate suggests AP2 is unreliable for individual patient decisions.
Area of Science:
- Critical Care Medicine
- Health Services Research
Background:
- The APACHE II (AP2) scoring system is widely used in clinical trials for risk-adjusted intensive care unit outcomes.
- While random errors may cancel out in large populations, concerns exist regarding AP2 accuracy for individual patient assessment.
- This study investigates AP2 scoring errors in a sepsis trial and their impact on individual patient classification.
Purpose of the Study:
- To report AP2 scoring errors made by researchers in a large sepsis trial.
- To model the consequences of these scoring errors for individual severe sepsis patients.
- To evaluate the appropriateness of AP2 as a sole decision-making tool for individual patient resource allocation.
Main Methods:
- Fifty-six researchers abstracted data from patient scenarios to calculate AP2 scores.
- Inter-observer reliability was assessed using Shrout-Fleiss methodology.
- Theoretical distribution curves modeled misclassification rates at an AP2 score cut-off of >=25, applied to the PROGRESS severe sepsis registry data.
Main Results:
- The overall AP2 score error rate was 86%, with individual variable errors ranging from 10% to 87%.
- Inter-observer reliability (intraclass correlation) was 0.51.
- Within the 17-28 AP2 score range (50% of PROGRESS registry patients), 70-85% of misclassified patients resided.
Conclusions:
- Individual patients are more likely to be scored incorrectly than correctly using AP2.
- Misclassification rates increase as the true AP2 score approaches the 25-point cut-off.
- APACHE II is an inappropriate sole tool for resource allocation decisions for individual patients.
More Related Videos
Related Concept Videos
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters assessment...
Kaplan-Meier Approach
Errors occurring during blood pressure monitoring
Several factors...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic illness...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
