Related Experiment Video
Updated: Jul 1, 2026

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)
Published on: August 28, 2021
Coefficients of Repeatability for Likely Change: Comparison between PROMIS Computer Adaptive Tests and Short Forms
M K Lee1, D Cella2, V Grzegorczyk3
1Department of Quantitative Health Sciences, Mayo Clinic, Rochester MN.
This study determined score changes needed to detect patient improvement using PROMIS computer adaptive testing (CAT) and short forms (SFs). Results show CAT and SFs offer comparable reliability for tracking patient change in health outcomes.
Area of Science:
- Health Outcomes Research
- Psychometrics
- Clinical Measurement
Background:
- Patient-reported outcomes (PROs) are crucial for assessing health status and treatment effectiveness.
- Computer Adaptive Testing (CAT) offers efficient and precise measurement of PROs.
- Short Forms (SFs) provide brief alternatives for PRO assessment, but their reliability compared to CAT needs clarification.
Purpose of the Study:
- To identify necessary score changes for detecting meaningful patient improvement in PROMIS CAT.
- To compare the reliability of PROMIS CAT with 4-item (SF 4a) and 8-item (SF 8a) short forms.
Main Methods:
- Utilized data from 5823 patients in a health system symptom management project.
- Calculated Coefficients of Repeatability (CRs) at 68% confidence for PROMIS CAT and SFs (SF 4a, SF 8a) across anxiety, depression, pain interference, and physical function domains.
- Compared CRs derived from item response theory standard errors with those from 90% confidence levels.
Main Results:
- CRs varied by domain and score range; for symptomatic ranges, CAT and SF 4a had CRs of |3|-|4|, while SF 8a had CRs of |2|-|3|.
- For scores within normal limits, CRs were higher, ranging from |3|-|8| for CAT, |3|-|14| for SF 4a, and |2|-|14| for SF 8a.
- Relaxing the confidence level from 90% to 68% reduced change thresholds significantly, especially for normal range scores.
Conclusions:
- PROMIS CAT and SFs demonstrate varying reliability depending on score range and domain.
- SF 8a showed slightly higher reliability than CAT in symptomatic ranges, while CAT was more reliable in healthier ranges.
- The choice of CAT or SFs should consider the specific patient population and desired precision for detecting change.
Related Concept Videos
Reliability and Validity
Comparing Experimental Results: Student's t-Test
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...

