Related Experiment Video
Updated: May 19, 2026

05:21
Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
Comparison of two Bayesian methods to detect mode effects between paper-based and computerized adaptive assessments:
1Department of Health Systems Science, College of Nursing, University of Illinois at Chicago, Chicago, IL 60612, USA. barthr@uic.edu
BMC Medical Research Methodology
|August 21, 2012
Summary
Detecting mode effects in computerized adaptive testing (CAT) is crucial. The modified robust Z (RZ) test offers better control of false positives than credible intervals (CrI) for detecting differential item functioning (DIF).
Area of Science:
- Psychometrics
- Health outcomes research
- Statistical modeling
Background:
- Computerized adaptive testing (CAT) is increasingly used for health outcome measures originally developed as paper-and-pencil (P&P) instruments.
- Differences in respondent item endorsement between CAT and P&P modes can introduce measurement error if not addressed.
- Accurate estimation of health outcomes requires identifying and correcting for mode effects.
Purpose of the Study:
- To propose and evaluate methods for detecting item-level mode effects in health outcome measures administered via CAT versus P&P.
- To assess the performance of a modified robust Z (RZ) test and 95% credible intervals (CrI) in identifying differential item functioning (DIF) due to mode of administration.
- To investigate the impact of various simulation conditions on the accuracy and power of these DIF detection methods.
Main Methods:
- Employed Bayesian estimation to derive posterior distributions of item parameters for detecting mode effects.
- Proposed two methods: a modified robust Z (RZ) test and 95% credible intervals (CrI) for item difficulty differences.
- Conducted a simulation study with varying data-generating models (1- & 2-parameter IRT), DIF sizes, percentage of DIF items, and mean trait level differences.
Main Results:
- Both RZ test and CrI demonstrated good to excellent control of false positives.
- The RZ test provided superior false positive control compared to CrI, particularly when items were easy to endorse or when mean trait levels differed between modes.
- Power to detect true positives was influenced by CAT item usage, item difficulty, and item discrimination, though overall detection power was suboptimal.
Conclusions:
- The RZ test is a robust method for controlling false positives in detecting mode-based DIF, outperforming CrI.
- While false positives were managed, the power to detect true DIF remains suboptimal, necessitating further research.
- Investigating the impact of prior assumptions and data conformity is crucial for refining these DIF detection methods, especially concerning easy-to-endorse items.
Related Concept Videos
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Comparing the Survival Analysis of Two or More Groups
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and Cox...
