相关实验视频
Updated: Jun 23, 2025

14:05
One Dimensional Turing-Like Handshake Test for Motor Intelligence
Published on: December 15, 2010
26.8K
人工智能透大学考试系统的现实世界测试:一个"图灵测试"案例研究
Peter Scarfe1, Kelly Watcham1, Alasdair Clarke2
1School Psychology and Clinical Language Sciences (PCLS), University of Reading, Reading, United Kingdom.
PloS one
|June 26, 2024
概括
大多数人工智能产生的学生工作在大学评估中没有被发现. 一项盲目的研究发现,94%的AI提交被遗漏,AI工作得分高于人类学生.
科学领域:
- 教育技术的教育技术
- 教育中的人工智能
- 学术诚信 学术诚信
背景情况:
- 像ChatGPT这样的先进人工智能工具的普及对教育评估完整性提出了重大挑战.
- 越来越多地依赖无监督评估,特别是自COVID-19大流行以来,增加了未被发现的AI辅助学术不当行为的风险.
- 确保学生工作的真实性对于保持学术资格的有效性至关重要.
研究的目的:
- 在真正的大学评估环境中调查100%人工智能生成提交的可检测性和性能.
- 评估未被发现的人工智能写作对高等教育学术诚信的影响.
主要方法:
- 进行了一项严格的盲目的研究,涉及将人工智能生成的工作提交到真实考试系统中.
- 人工智能提交的作品占英国大学所有学年五个本科心理学模块的100%工作.
- 评估是在标记者没有事先知道提交内容是人工智能生成的的情况下进行的.
主要成果:
- 显著的94%的人工智能生成的提交没有被考试系统和标记器检测到.
- 人工智能提交的作品获得的成绩平均比人类学生获得的成绩高出一半.
- 有83.4%的概率,人工智能提交的内容将超过同一个模块内的真实学生提交的随机选择.
结论:
- 目前的评估实践对未被检测到的AI产生的内容非常脆弱,对学术完整性构成重大威胁.
- 人工智能写作工具可以实现比平均学生工作更高的学术表现,需要重新评估评估策略.
- 大学必须开发强大的方法来检测人工智能产生的工作,并调整评估设计,以减轻与教育中人工智能相关的风险.
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
The Stanford Prison Experiment
23.1K
The famous and controversial Stanford Prison Experiment, conducted by social psychologist Philip Zimbardo and his colleagues at Stanford University, demonstrated the power of social roles, social norms, and scripts.
23.1K
Triarchic Theory of Intelligence
7.8K
Robert Sternberg's triarchic theory of intelligence posits that intelligence is composed of three distinct but interrelated components: analytical, creative, and practical intelligence.
7.8K
Measures of Intelligence
7.2K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.2K

