Related Experiment Video
Updated: Aug 8, 2026

The Participant-Reported Implementation Update and Score (PRIUS): A Novel Method for Capturing Implementation-Related Data Over Time
Published on: February 19, 2021
Interobserver variability in data collection of the APACHE II score in teaching and community hospitals
L M Chen1, C M Martin, T L Morrison
1Critical Care Research Network, London Health Sciences Centre, Ontario, Canada.
Objectives:
To examine interobserver reliability of the Acute Physiologic and Chronic Health Evaluation (APACHE) II score and identify major causes of variability in data collection.
Design:
Descriptive, comparative analysis.
Setting:
Nine intensive care units in two teaching and six community hospitals
Subjects:
A random sample of 342 patient records selected from a network database.
Intervention:
None.
Measurements And Main Results:
Data were reabstracted and compared with the original records. Individual physiologic points derived from the APACHE II scoring system (instead of the actual physiologic values) were compared using the kappa statistic. Paired measurements of the continuous variables were compared using the interclass correlation coefficient and Bland-Altman plots. Excellent agreement was found in most demographic, admission, and discharge data. The system failure requiring intensive care unit admission was consistently identified by both data collectors in 88% of cases, but only 66% agreed on the exact admitting diagnosis. For APACHE II score components, the kappa statistic ranged from 0.315 for the Glasgow Coma Scale point to 0.976 for the age point. Significant disagreement regarding the probability of death derived from the APACHE II model was evident in some patient records. Overall agreement among groups of patients regarding the APACHE II score was good, however, with no significant difference in the mean score (20.2 vs. 20.1; p = .758). The predicted mortality from the reabstracted data was 30%, similar to the 27% predicted mortality from the original data (p = .380).
Conclusion:
Reliability of data collection varied widely in different components of the APACHE II probability-of-death model. Significant discrepancies in some components suggested a lack of explicit definitions and timing for consistent data collection between institutions or between data collectors. Nonetheless, variability resulting from data collection appears to be randomly distributed, so that comparisons of group means are valid.
Related Concept Videos
Surveys
Reliability and Validity
SBAR I: Understanding the Concept
Standardized methods of communication have been developed to ensure that information is...
SBAR II: Application of SBAR
SBAR Report from a Nurse to a Health Care Provider
S: "Hello, Dr. Smith. This is Jane, RN, from the Med Surg unit. I am calling to tell you about Ms. White in Room 210, who is experiencing increased pain and redness at her incision site. Her recent...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...