通过自动化写作评估程序,提高基于书面表达的课程测量可行性
Michael Matta1, Milena A Keller-Margulis1, Sterett H Mercer2
1Department of Psychological, Health, and Learning Sciences, University of Houston.
School psychology (Washington, D.C.)
|March 20, 2025
概括
自动化写作评估 (AWE) 工具准确地评分学生的写作,与更简单的指标的人类评分密切匹配. 这些自动化得分显示没有偏见,并有效地预测国家写作测试的未来表现.
科学领域:
- 教育技术的教育技术
- 心理测量 心理测量 心理测量
- 写作评估 写作评估
背景情况:
- 自动写作评估 (AWE) 程序为学生的写作提供了可行的替代方案.
- 现有的AWE工具需要严格验证准确性,预测有效性和潜在偏差.
研究的目的:
- 评估基于书面表达的课程测量 (WE-CBM) 的自动成绩的准确性,预测有效性和偏差.
- 将自动化WE-CBM指标与传统的手工计分指标进行比较.
主要方法:
- 采集了2-5年级的722名学生的写作样本,使用3分钟的WE-CBM任务.
- 手工评分了四个WE-CBM指标,并使用基于计算机的方法为相同的指标生成了自动评分.
- 将自动化得分与手工得分的指标进行比较,并分析了对国家强制性写作测试绩效的预测.
主要成果:
- 简单的自动化指标 (词汇总数,正确拼写的词汇) 与手工计算的分数非常接近.
- 在更复杂的指标 (正确的单词序列) 中发现了小小的差异.
- 自动化得分准确地预测了州考试成绩,类似于手工得分的指标,没有针对非洲裔美国人和西班牙裔学生的偏见.
结论:
- 自动化WE-CBM评分表现出高准确性和预测有效性,与传统方法相比.
- 自动评分显示没有证据表明少数族裔学生群体的偏见.
- 调查结果支持在书面评估中用于教育决策的自动分数的使用.
更多相关视频
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
649
14:43Universal Screening for Prevention of Reading, Writing, and Math Disabilities in Spanish
Published on: July 18, 2020
7.9K
相关概念视频
Binet's Contribution to Measures of Intelligence
1.2K
Alfred Binet, along with his student Théophile Simon, was tasked by the French Ministry of Education in 1904 to create a method for identifying students who struggled to learn through conventional classroom instruction. This initiative aimed to address overcrowding by placing such students in specialized schools. Binet and Simon developed an intelligence test comprising 30 tasks, ranging from simple commands, like touching one's nose or ear, to more complex tasks, such as drawing...
1.2K
Reliability and Validity
12.6K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.6K
Wechsler's Contribution to Measures of Intelligence
1.4K
David Wechsler, a psychologist who worked with World War I veterans, developed a significant IQ test in 1939 called the Wechsler-Bellevue Intelligence Scale. This test was innovative because it combined several subtests that measured both verbal and nonverbal skills, reflecting Wechsler's belief that intelligence is a global capacity involving purposeful action, rational thinking, and effective interaction with the environment. This test later evolved into the Wechsler Adult Intelligence...
1.4K
Measures of Intelligence
5.9K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
5.9K
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
