对于PISA的高度适应性测试设计的方法方面.
Aron Fink1, Christoph König1, Andreas Frey1
1Institute of Psychology, Goethe University Frankfurt, Frankfurt, Germany.
Frontiers in psychology
|October 2, 2024
概括
高度适应性测试设计 (HAT) 为PISA等评估最大限度地提高了项目选择适应性. 这种方法使用计算机算法来平衡适应性和测试约束,改善学生的体验.
科学领域:
- 教育测量的教育测量.
- 计算机化的适应性测试
- 心理测量 心理测量 心理测量
背景情况:
- 国际学生评估计划 (PISA) 需要高效和适应性测试方法.
- 现有的计算机自适应测试 (CAT) 方法在处理复杂的约束和嵌套项目结构时存在局限性.
研究的目的:
- 描述高度适应性测试 (HAT) 设计的方法论和统计基础.
- 提出一种新的算法,在评估约束范围内最大限度地提高项目选择的适应性.
主要方法:
- 使用R编程语言开发了一个HAT算法.
- 综合已建立的CAT方法来解决嵌套项目,维度相关性,约束管理和项目位置效应.
- 专注于增强学生的考试体验.
主要成果:
- HAT设计允许在项目选择中实现最大的适应性.
- 该算法有效地管理PISA的特定约束.
- 该方法改进了标准的CAT方法,通过结合多个因素来实现最佳的测试构建.
结论:
- 在大规模评估中,HAT设计为适应性测试提供了一个强大的框架.
- 提供的 R 代码有助于实现和适应 HAT,以便在未来进行研究和评估.
- 这种方法可以激发新的适应性测试设计的开发,以平衡适应性与实际约束.
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Decision Making: P-value Method
5.3K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.3K
Accuracy and Errors in Hypothesis Testing
180
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
180
Binet's Contribution to Measures of Intelligence
1.2K
Alfred Binet, along with his student Théophile Simon, was tasked by the French Ministry of Education in 1904 to create a method for identifying students who struggled to learn through conventional classroom instruction. This initiative aimed to address overcrowding by placing such students in specialized schools. Binet and Simon developed an intelligence test comprising 30 tasks, ranging from simple commands, like touching one's nose or ear, to more complex tasks, such as drawing...
1.2K
Study Design in Statistics
7.9K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
7.9K
Group Design
8.9K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
8.9K


