並列テスト設定における最適キャリブレーション設計の開発:達成度テストにおける混合形式項目の効率向上
Frank Miller1,2, Ellinor Fackle-Fornius1
1Stockholm University.
Psychometrika
|February 25, 2026
まとめ
本研究は、並列テストのための最適キャリブレーション設計を導入し、大規模達成度テストにおける項目キャリブレーションの効率を向上させます。この手法は、応答時間が異なる混合形式テストの精度を向上させます。
科学分野:
- 教育測定
- 心理測定学
- テスト理論
背景:
- 大規模達成度テストでは、運用上の使用のために定期的な項目キャリブレーションが必要です。
- 既存の最適割り当て手法は、主に逐次的な受験者到着を扱っています。
- キャリブレーション設定では、同時または並列のテスト管理が一般的です。
研究 の 目的:
- 並列テスト設定のための最適キャリブレーション設計を開発すること。
- 提案手法の効率向上を調査すること。
- 実際のキャリブレーションシナリオにおける手法の適用可能性を実証すること。
主な方法:
- 並列テスト管理に合わせた最適キャリブレーション設計を開発しました。
- この手法は混合形式項目を扱い、応答時間のばらつきを考慮します。
- スウェーデン国家数学テストの項目をキャリブレーションするためにこの手法を適用しました。
主要な成果:
- 提案手法はキャリブレーション効率を大幅に向上させます。
- 実際のケーススタディで実装に成功したことを実証しました。
- この設計は、応答時間のばらつきが大きい混合形式テストに有効です。
結論:
- 新しい最適キャリブレーション設計は、並列テスト設定において効率的かつ実用的です。
- この手法は、大規模評価における項目のキャリブレーションに大幅な改善を提供します。
- このアプローチは、教育測定における実践的な課題に対処する混合形式テストに適しています。
さらに関連する動画
10:26Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
4.5K
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
1.3K
関連する概念動画
Goodness-of-Fit Test
9.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
9.3K
Expected Frequencies in Goodness-of-Fit Tests
8.7K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.7K
One-Way ANOVA: Equal Sample Sizes
4.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.2K
Friedman Two-way Analysis of Variance by Ranks
522
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
522
One-Way ANOVA: Unequal Sample Sizes
6.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
6.8K
