Related Experiment Video
Updated: Sep 25, 2026

An Experimental Paradigm for the Prediction of Post-Operative Pain (PPOP)
Published on: January 27, 2010
Reliability, construct validity, responsiveness, and minimal important difference of the Numerical Rating Scale for
Alexandre Delgado1, Melania M Amorim1, Guilherme Tavares de Arruda2
1Instituto de Medicina Integral Prof. Fernando Figueira (IMIP), Recife, PE, Brazil.
Objective:
To evaluate the measurement properties of the Numerical Rating Scale for assessing pain severity during labor, including test-retest reliability, measurement error, construct validity, responsiveness, and minimum important difference (MID).
Methods:
This longitudinal study was a secondary analysis of a randomized clinical trial involving 200 usual-risk pregnant women in active labor. Participants were randomized to either a pelvic biomechanics intervention protocol using kinesiotherapy with a Swiss ball or usual care. Pain severity was assessed using the 11-point NRS at baseline and after 30, 60, and 90 min. Test-retest reliability was evaluated using the intraclass correlation coefficient (ICCagreement). Measurement error was analyzed using the standard error of measurement (SEMagreement), smallest detectable change (SDCagreement), and limits of agreement (LoA). Construct validity was assessed by comparing pain scores between nulliparous and multiparous women. Responsiveness was evaluated through mixed linear models and standardized effect sizes.
Results:
The NRS demonstrated sufficient reliability, with ICCagreement values ranging from 0.713 to 0.963 across assessment intervals. Measurement error progressively decreased over time, with SEMagreement values ranging from 1.03 to 0.23 and SDCagreement values ranging from 2.86 to 0.64. Construct validity was supported, as nulliparous women consistently reported higher pain severity than multiparous women at baseline and follow-up assessments. Responsiveness analysis revealed significant group-by-time interaction effects, indicating distinct pain trajectories between groups. The intervention group showed significantly lower pain scores than the control group at 30, 60, and 90 min (p < 0.001). Between-group differences exceeded the proposed MID threshold of two points at 30 and 60 min. A large effect size was observed for responsiveness (Cohen's d = 1.17).
Conclusion:
The NRS demonstrated sufficient reliability, construct validity, and interpretability, and partially sufficient responsiveness, for assessing labor pain severity. The instrument was capable of discriminating between clinically distinct groups and detecting meaningful changes over time during labor. These findings support the use of the NRS as a practical and clinically relevant tool for labor pain assessment in both research and clinical settings.
