Related Experiment Video
Updated: Feb 25, 2026

Use of a Video Scoring Anchor for Rapid Serial Assessment of Social Communication in Toddlers
Published on: March 14, 2018
Inter-rater reliability of two paediatric early warning score tools
Claus S Jensen1,2,3, Hanne Aagaard4, Hanne V Olesen4
1Department of Paediatrics and Adolescent Medicine, Herlev Gentofte Hospital.
Insights
Nurses demonstrated good to very good reliability when using two paediatric early warning score (PEWS) systems. This indicates PEWS tools are dependable for detecting patient deterioration in clinical settings.
Area of Science:
- Pediatric critical care
- Healthcare quality improvement
- Clinical assessment tools
Background:
- Paediatric early warning score (PEWS) assessment tools aid in detecting subtle changes indicating clinical deterioration.
- The reliability of PEWS tools depends on the accuracy of caregiver data collection and documentation.
Purpose of the Study:
- To evaluate the inter-rater reliability of PEWS systems among nurses.
- Assessing the consistency of PEWS assessments between different nurses.
Main Methods:
- Study conducted in five pediatric departments in the Central Denmark Region.
- Inter-rater reliability assessed via parallel observations with 108 children and 69 nurses.
- Statistical analysis included Intraclass Correlation Coefficient, Fleiss' κ, and Bland-Altman limits of agreement.
Main Results:
- Intraclass correlation coefficients for aggregated PEWS scores were high (0.98 and 0.95).
- Fleiss' κ values for individual measurements indicated good to very good agreement (0.70–1.0).
- Nurses assigned identical aggregated scores in 76% of cases, with 98% of scores differing by ≤1 point.
Conclusions:
- The study demonstrated good to very good inter-rater reliability for the two PEWS models evaluated.
- Findings support the consistent application of these PEWS tools by nursing staff in the Central Denmark Region.
Background:
Paediatric early warning score (PEWS) assessment tools can assist healthcare providers in the timely detection and recognition of subtle patient condition changes signalling clinical deterioration. However, PEWS tools instrument data are only as reliable and accurate as the caregivers who obtain and document the parameters.
Objective:
The aim of this study is to evaluate inter-rater reliability among nurses using PEWS systems.
Design:
The study was carried out in five paediatrics departments in the Central Denmark Region. Inter-rater reliability was investigated through parallel observations. A total of 108 children and 69 nurses participated. Two nurses simultaneously performed a PEWS assessment on the same patient. Before the assessment, the two participating nurses drew lots to decide who would be the active observer. Intraclass correlation coefficient, Fleiss' κ and Bland-Altman limits of agreement were used to determine inter-rater reliability.
Results:
The intraclass correlation coefficients for the aggregated PEWS score of the two PEWS models were 0.98 and 0.95, respectively. The κ value on the individual PEWS measurements ranged from 0.70 to 1.0, indicating good to very good agreement. The nurses assigned the exact same aggregated score for both PEWS models in 76% of the cases. In 98% of the PEWS assessments, the aggregated PEWS scores assigned by the nurses were equal to or below 1 point in both models.
Conclusion:
The study showed good to very good inter-rater reliability in the two PEWS models used in the Central Denmark Region.

