Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reliability and Validity01:29

Reliability and Validity

Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
Comparing Experimental Results: Student's t-Test01:09

Comparing Experimental Results: Student's t-Test

The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
Testing a Claim about Standard Deviation01:19

Testing a Claim about Standard Deviation

A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Bonferroni Test01:10

Bonferroni Test

The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same authorSame journal

Crosswalk between PROMIS computer adaptive tests and numerical rating scales in cancer patients: Anxiety, depression, pain interference, physical function.

Advances in patient-reported outcomes·2026
Same author

Development of the novel nontuberculous mycobacterial pulmonary disease symptoms scale (NTM-SS).

Quality of life research : an international journal of quality of life aspects of treatment, care and rehabilitation·2026
Same author

United Kingdom value set for the functional assessment of cancer therapy eight dimension (FACT-8D) preference-based quality of life instrument.

The European journal of health economics : HEPAC : health economics in prevention and care·2025
Same author

Health-related quality-of-life profile and clinical outcomes in first-line advanced renal cell carcinoma: a modeling analysis based on the CheckMate 9ER study.

ESMO open·2025
Same author

Derivation of a simple risk calculator for predicting clinical worsening in patients with pulmonary hypertension due to interstitial lung disease.

JHLT open·2025
Same author

Characteristics and clinical outcomes of breast cancer in young BRCA carriers according to tumor histology.

ESMO open·2024

Related Experiment Video

Updated: Jul 1, 2026

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)
10:58

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)

Published on: August 28, 2021

Coefficients of Repeatability for Likely Change: Comparison between PROMIS Computer Adaptive Tests and Short Forms.

M K Lee1, D Cella2, V Grzegorczyk3

  • 1Department of Quantitative Health Sciences, Mayo Clinic, Rochester MN.

Advances in Patient-Reported Outcomes
|June 30, 2026
PubMed
Summary

This study determined score changes needed to detect patient improvement using PROMIS computer adaptive testing (CAT) and short forms (SFs). Results show CAT and SFs offer comparable reliability for tracking patient change in health outcomes.

Keywords:
Coefficient of repeatabilityComputer adaptive testingItem response theoryLikely change indexPROMISReliable change indexShort formsWithin-individual meaningful change

More Related Videos

Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

Advancing Dyslexia Assessment in Children Through Computerized Testing
09:00

Advancing Dyslexia Assessment in Children Through Computerized Testing

Published on: August 16, 2024

Related Experiment Videos

Last Updated: Jul 1, 2026

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)
10:58

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)

Published on: August 28, 2021

Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

Advancing Dyslexia Assessment in Children Through Computerized Testing
09:00

Advancing Dyslexia Assessment in Children Through Computerized Testing

Published on: August 16, 2024

Area of Science:

  • Health Outcomes Research
  • Psychometrics
  • Clinical Measurement

Background:

  • Patient-reported outcomes (PROs) are crucial for assessing health status and treatment effectiveness.
  • Computer Adaptive Testing (CAT) offers efficient and precise measurement of PROs.
  • Short Forms (SFs) provide brief alternatives for PRO assessment, but their reliability compared to CAT needs clarification.

Purpose of the Study:

  • To identify necessary score changes for detecting meaningful patient improvement in PROMIS CAT.
  • To compare the reliability of PROMIS CAT with 4-item (SF 4a) and 8-item (SF 8a) short forms.

Main Methods:

  • Utilized data from 5823 patients in a health system symptom management project.
  • Calculated Coefficients of Repeatability (CRs) at 68% confidence for PROMIS CAT and SFs (SF 4a, SF 8a) across anxiety, depression, pain interference, and physical function domains.
  • Compared CRs derived from item response theory standard errors with those from 90% confidence levels.

Main Results:

  • CRs varied by domain and score range; for symptomatic ranges, CAT and SF 4a had CRs of |3|-|4|, while SF 8a had CRs of |2|-|3|.
  • For scores within normal limits, CRs were higher, ranging from |3|-|8| for CAT, |3|-|14| for SF 4a, and |2|-|14| for SF 8a.
  • Relaxing the confidence level from 90% to 68% reduced change thresholds significantly, especially for normal range scores.

Conclusions:

  • PROMIS CAT and SFs demonstrate varying reliability depending on score range and domain.
  • SF 8a showed slightly higher reliability than CAT in symptomatic ranges, while CAT was more reliable in healthier ranges.
  • The choice of CAT or SFs should consider the specific patient population and desired precision for detecting change.