Related Experiment Video
Updated: Aug 8, 2025

07:13
A Two-interval Forced-choice Task for Multisensory Comparisons
Published on: November 9, 2018
11.0K
Multidimensional Forced-Choice CAT With Dominance Items: An Empirical Comparison With Optimal Static Testing Under
Yin Lin1,2, Anna Brown1, Paul Williams3
1University of Kent, Canterbury, UK.
Educational and Psychological Measurement
|March 3, 2023
Summary
This study explored forced-choice computerized adaptive tests (FC CATs) using dominance items. While adaptive item selection improved precision, FC CATs showed no advantage over static tests at shorter lengths.
Area of Science:
- Organizational Psychology
- Psychometrics
Background:
- Forced-choice computerized adaptive tests (FC CATs) often use ideal-point items.
- Research on FC CATs utilizing dominance items, despite their historical prevalence, is limited and largely based on simulations.
Purpose of the Study:
- To empirically investigate a FC CAT employing dominance items based on the Thurstonian Item Response Theory model.
- To examine the impact of adaptive item selection and social desirability balancing on score distributions, measurement accuracy, and participant perceptions.
- To compare the adaptive FC CAT against nonadaptive optimal tests to assess the return on investment for adaptive assessments.
Main Methods:
- An empirical study was conducted using research participants.
- A FC CAT with dominance items (Thurstonian IRT model) was trialed.
- Nonadaptive optimal tests were used as a baseline for comparison.
Main Results:
- Adaptive item selection was confirmed to enhance measurement precision.
- FC CATs did not demonstrate a significant advantage over optimal static tests at shorter test lengths.
- The study analyzed practical issues including score distributions, accuracy, and participant perceptions.
Conclusions:
- Adaptive item selection offers benefits for measurement precision in FC CATs.
- The operational and psychometric advantages of FC CATs over optimized static tests are most apparent at longer test lengths.
- Implications for designing and deploying FC assessments in research and practice were discussed, considering both psychometric and operational factors.
Related Concept Videos
Multiple Comparison Tests
4.0K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.0K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Friedman Two-way Analysis of Variance by Ranks
268
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
268
McNemar's Test
342
McNemar's Test is a nonparametric statistical test used to determine if there is a significant difference in proportions between two related groups when the outcome is binary (e.g., yes/no, success/failure). It is beneficial when we have paired data, such as pre-test/post-test designs, where the same subjects are measured under two different conditions. The test is named after the statistician Quinn McNemar, who introduced it in 1947. It is commonly used in situations where subjects are...
342
Goodness-of-Fit Test
3.7K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.7K
Wilcoxon Signed-Ranks Test for Matched Pairs
187
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
187

