概括
新的统计评估计算机自适应测试 (CATs) 通过比较条件标准测量误差 (CSEMs) 来评估得分可比性. 该研究发现,CAT通常具有很高的得分可比性,得分极端的差异很大.
科学领域:
- 教育测量和统计学
- 心理测量 心理测量 心理测量
- 计算机化的适应性测试 (CAT)
背景情况:
- 计算机化适应性测试 (CAT) 在教育评估中被广泛使用.
- 确保不同CAT表格的分数可比性对于有效的分数解释至关重要.
- 评估得分可比性的现有方法存在局限性.
研究的目的:
- 引入两个新的统计数据来量化CAT的得分可比性.
- 使用这些新统计数据,评估使用这些新统计数据的替代CAT表格的得分可比性.
- 评估得分范围和样本大小对得分可比性的影响.
主要方法:
- 开发了两个基于比较具有相同等级分数的受试者测量条件标准误差 (CSEM) 的统计数据.
- 一个统计数据评估了个人规模得分的可比性;另一个统计数据评估了整体得分的可比性.
- 将统计数据应用于3-8年级的阅读和数学CAT数据.
主要成果:
- 两个CAT都表现出相当高的得分可比性.
- 在得分范围非常高或非常低时,得分可比性下降.
- 使用最低20%的得分者有时会比使用所有学生产生更高的整体得分可比性.
结论:
- 拟议的统计数据有效地衡量了CAT分数的可比性.
- CATs通常表现出良好的分数可比性,尽管它因分数级别而异.
- 这些发现为开发和评估CAT提供了有价值的见解.
更多相关视频
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
792
10:58Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics BM-PROMA
Published on: August 28, 2021
4.5K
相关概念视频
Measures of Intelligence
7.3K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.3K
Review and Preview
7.6K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
7.6K
Comparing Experimental Results: Student's t-Test
1.6K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
1.6K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
1.6K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.6K
