检测人工智能对在线问卷的腐败行为
Benjamin Lebrun1, Sharon Temtsin2, Andrew Vonasch1
1School of Psychology, Speech, and Hearing, University of Canterbury, Christchurch, New Zealand.
Frontiers in robotics and AI
|February 19, 2024
概括
人工智能 (AI) 可以为在线研究生成文本,但人类检测准确率只有76%. 目前的AI检测系统无法使用,威胁到在线调查数据质量.
科学领域:
- 社会科学 社会科学 社会科学
- 计算机科学 计算机科学
- 数据科学数据科学数据科学
背景情况:
- 在线调查问卷利用众包来提高效率和成本效益.
- 人工智能 (AI) 的进步,特别是大型语言模型 (LLM),使自动填写表格和生成文本成为可能.
- 这对通过在线调查收集的数据的完整性构成重大威胁.
研究的目的:
- 通过人类评估员和自动化AI检测系统在线研究中评估AI生成文本的可检测性.
- 在面对人工智能驱动的欺诈性提交时,评估当前方法在确保数据质量的有效性.
主要方法:
- 人类参与者被要求在在线研究环境中区分人写和人工智能生成的文本.
- 评估了自动化AI检测系统,以确定它们识别AI产生的响应的能力.
- 记录和分析了人类和自动检测的准确率.
主要成果:
- 人类评估人员在识别人工智能生成的文本时达到76%的准确率,这超出了机会,但不足以确保高质量的数据.
- 目前的自动AI检测系统在检测AI生成的提交时被证明是完全无效的.
- 该研究强调了当前在线研究中维护数据完整性的方法论中的关键差距.
结论:
- 仅仅依靠人类的注意力检查已经变得不足以保证数据质量,因为人工智能产生的内容越来越复杂.
- 众包平台必须开发系统解决方案,以解决人工智能驱动的数据制造问题,因为当前的检测工具不足.
- 越来越多的人工智能提交的风险使欺诈检测的成本变得过高,破坏了在线问卷的实用性.
相关概念视频
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Surveys
14.8K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.8K
Ethics in Research
23.0K
Today, scientists agree that good research is ethical in nature and is guided by a basic respect for human dignity and safety. However, this has not always been the case. Modern researchers must demonstrate that the research they perform is ethically sound.
23.0K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K


