相关实验视频
Updated: Jun 28, 2025

12:10
Assessment of Mouse Judgment Bias through an Olfactory Digging Task
Published on: March 4, 2022
2.6K
人类标记的机器常识推理基准的噪音审计
Mayank Kejriwal1, Henrique Santos2, Ke Shen3
1Information Sciences Institute, University of Southern California, Marina del Rey, 90292, USA. kejriwal@isi.edu.
Scientific reports
|April 13, 2024
概括
人类对人工智能基准的判断是杂的. 这项研究揭示了机器常识推理数据集中的大量噪音,影响了人类和像ChatGPT这样的AI系统的性能估计.
科学领域:
- 人工智能的人工智能
- 认知科学 认知科学
- 人与计算机的交互
背景情况:
- 评估大型语言模型 (LLM) 对人工智能进步至关重要.
- 传统的AI基准测试依赖于一个单一的"基础真理"进行比较.
- 心理学研究表明,人类的判断任务往往包含大量的噪音.
研究的目的:
- 在机器常识推理中对人类标记的基准进行详细的噪音审计.
- 在受控,高质量的标签和现实的众包条件下分析噪音.
- 评估噪音对人工智能系统性能估计和人类判断的影响.
主要方法:
- 应用卡尼曼的框架来审核人类标记的常识推理基准中的噪音.
- 在两个环境中进行审计:小规模,高质量的标签和大规模众包.
- 标签过程中的量化水平,模式和系统噪声.
主要成果:
- 连在高质量设置中也始终发现不小的水平,模式和系统噪音量.
- 众包环境中的噪音水平与更高质量的环境相美.
- 噪音显著影响了性能估计:人类系统高达10%,ChatGPT超过4%.
结论:
- 在AI基准测试中假设单一的"基础真理"可能是有缺陷的,特别是对于需要人类判断的任务.
- 标记噪声可以大大改变感知到的系统性能,需要对方法进行重新评估.
- 研究结果强调了人工智能研究中需要强大的噪音意识评估策略.
更多相关视频
相关概念视频
Heuristics
89
Heuristics are problem-solving strategies that use mental shortcuts to simplify decision-making. Unlike algorithms, which must be followed precisely to achieve a correct result, heuristics offer a general problem-solving framework. They save time and energy but can sometimes lead to less rational decisions.
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
89
Deductive Reasoning
55.3K
Deductive reasoning, or deduction, is the type of logic used in hypothesis-based science. In deductive reasoning, the pattern of thinking moves in the opposite direction as compared to inductive reasoning, which means that it uses a general principle or law to predict specific results. From those general principles, a scientist can deduce and predict the specific results that would be valid as long as the general principles are valid.
For example, a researcher can deduce specific predictions...
For example, a researcher can deduce specific predictions...
55.3K
Inductive Reasoning
60.4K
Inductive reasoning is a form of logical thinking that uses related observations to arrive at a general conclusion. It is uncertain and operates in degrees to which the conclusions are credible. As such, inductive arguments can be weak or strong, rather than valid or invalid, and conclusions can be used to formulate testable, falsifiable hypotheses.
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
60.4K
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Types of Errors: Detection and Minimization
1.6K
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
1.6K
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K

