人工智能辅助评分和个性化反在大型政治学课堂:随机对照试验的结果
Tobias Heinrich1, Spencer Baily2, Kuan-Wu Chen3
1Department of Political Science, The University of Houston, Houston, Texas, United States of America.
PloS one
|August 19, 2025
概括
大型语言模型 (LLM) 的帮助可以显著提高教师的生产力,在大班级中对短答案问题进行评分. 人工智能辅助的评分有效地模仿了人类的评分,为学生保留了批判性思维技能发展的机会.
科学领域:
- 教育技术的教育技术
- 教育中的人工智能
- 高等教育教学的教学.
背景情况:
- 对教练来说,对短答题进行评分是耗时的.
- 激励有利于多选项测试,限制了批判性思维技能的发展.
- 大型语言模型 (LLM) 提供了对辅助的分级的潜力.
研究的目的:
- 评估AI辅助评分对短答题问题的有效性.
- 将人工智能辅助的评分与传统的人类评分方法进行比较.
- 评估LLM辅助对讲师生产力和学生反的影响.
主要方法:
- 在四个本科课程中进行了一项随机对照试验.
- 该研究在2023/2024学年期间涉及近300名学生.
- 人工智能辅助的评分与短答案评估的人类评分进行了比较.
主要成果:
- 人工智能辅助分级显示了与人类分级相似的结果.
- 提高了指导员对短答题进行评分的生产力.
- 人工智能辅助评分成功模仿了小班环境中的教师评分.
结论:
- 通过LLM辅助评分是提高大型本科课程效率的可行工具.
- 人工智能工具可以支持教师提供个性化的反,而不会牺牲评估质量.
- 将人工智能集成到评分过程中可以帮助保持发展批判性思维技能的机会.
更多相关视频
10:26Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
4.1K
10:17Improving Student Outcomes with an Adaptable Molecular Cloning Course-Based Undergraduate Research Experience
Published on: November 15, 2024
1.2K
相关概念视频
Surveys
15.4K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
15.4K
Group Design
9.6K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
9.6K
Randomized Experiments
7.2K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
7.2K
Social Loafing
35.9K
Another way in which a group presence can affect performance is social loafing—the exertion of less effort by a person working together with a group. Social loafing occurs when our individual performance cannot be evaluated separately from the group. Thus, group performance declines on easy tasks (Karau & Williams, 1993). Essentially individual group members loaf and let other group members pick up the slack. Because each individual’s efforts cannot be evaluated,...
35.9K
Reliability and Validity
13.2K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.2K
Self-Evaluation: Self-Enhancement and Self-Verification
5.3K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
5.3K
