Planning a Study for Testing the Rasch Model given Missing Values due to the use of Test-booklets
Takuya Yanagida1, Klaus D Kubinger, Dieter Rasch
1Takuya Yanagida, School of Applied Health and Social Studies, University of Applied Sciences Upper Austria, Garnisonstrabe 21, 4020 Linz, Austria, Takuya.Yanagida@fh-linz.at.
Summary
This study validates a sample size approach for Rasch model testing, even with missing data common in large-scale assessments. The method remains effective when using multiple test booklets, ensuring reliable Type-I risk and test power.
Area of Science:
- Psychometrics
- Educational Measurement
- Statistical Modeling
Background:
- The Rasch model is widely used for achievement test calibration in psychology and education.
- Existing sample size determination methods often lack statistical rigor for data sampling.
- Previous simulation studies by Kubinger et al. (2009, 2011) focused on complete data scenarios.
Purpose of the Study:
- To investigate the applicability of a sample size determination approach for the Rasch model in the presence of missing data.
- To evaluate if the approach by Kubinger et al. (2009, 2011) remains valid when using multiple test booklets in large-scale assessments.
- To assess the impact of designed missing values on Type-I risk and test power.
Main Methods:
- A simulation study was conducted to analyze the three-way analysis of variance design with mixed classification.
- The study specifically addressed scenarios with missing values due to the use of multiple test booklets.
- The performance of the sample size approach was evaluated under conditions with dichotomous item response data.
Main Results:
- The simulation results indicate that the sample size determination approach is effective even with designed missing values.
- The use of test booklets did not significantly alter the actual Type-I risk or the power of the Rasch model test.
- Effectiveness was observed when examinees provided a comparable amount of information per item.
Conclusions:
- The validated approach for sample size determination is robust to missing data patterns typical in large-scale assessments.
- Researchers can confidently apply this method when using multiple test booklets, provided sufficient examinee information per item.
- It is recommended that researchers conduct their own simulations for specific scenarios due to the study's focused scope.
More Related Videos
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
4.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.5K
Expected Frequencies in Goodness-of-Fit Tests
8.8K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.8K
Errors In Hypothesis Tests
6.2K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
6.2K
Introduction to Test of Independence
3.1K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
3.1K
Censoring Survival Data
636
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
636
Goodness-of-Fit Test
9.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
9.4K


