可信度机器常识 (PMC) 数据集:用于研究大型语言模型中的可信度的大量众包人类注释的数据集
Navapat Nananukul1, Ke Shen1, Mayank Kejriwal1
1University of Southern California, 4676 Admiralty Way, Suite 1001 Marina del Rey, CA 90292, USA.
Data in brief
|September 19, 2024
概括
本研究引入了用于常识推理的新数据集,重点是可信度评估. 它将人类注释与人工智能模型输出进行比较,以推进人工智能能力.
科学领域:
- 人工智能的人工智能
- 认知科学 认知科学
- 自然语言处理自然语言处理.
背景情况:
- 常识推理是AI的一个关键挑战.
- 可信度评估,确定陈述的可能性,是一个未被充分探索的领域.
- 微妙的,多标签的人类注释对于开发AI模型至关重要.
研究的目的:
- 创建一个全面的数据集,用于常识可信度评估.
- 将人类的可信性判断与高级人工智能模型 (GPT-3.5,GPT-4) 的判断进行比较.
- 为培训和评估人工智能的资源提供一种类似人类的常识推理.
主要方法:
- 使用众包 (亚马逊机械土耳其) 重新注释现有数据集 (SemEval-2020任务4).
- 在2000个句子中收集1万个独特的注释,每个句子有5个注释.
- 用句子促使GPT-3.5和GPT-4模型,并将它们的可信性标签与人类共识进行比较.
主要成果:
- 一个新的数据集 (PMC-Dataset) 产生了细微的标签 (可信,不可信,模糊).
- 这项研究为分析人类与机器之间的比较提供了一个基准,对合理性的常识推理进行了分析.
- 最初的比较表明,人工智能模型有可能与人类的判断保持一致.
结论:
- 该PMC-数据集是人工智能研究在常识推理中的宝贵资源.
- 它促进了人工智能模型的开发,可以更好地理解和复制人类可信度评估.
- 这项工作支持人工智能应用的进步,需要对日常事件和陈述有细微的理解.
相关概念视频
Hypothesis: Accept or Fail to Reject?
27.6K
The outcome of any hypothesis testing leads to rejecting or not rejecting the null hypothesis. This decision is taken based on the analysis of the data, an appropriate test statistic, an appropriate confidence level, the critical values, and P-values. However, when the evidence suggests that the null hypothesis cannot be rejected, is it right to say, 'Accept' the null hypothesis?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
27.6K
Modeling and Similitude
249
Scaled modeling is a fundamental technique in engineering, enabling the study of large and complex systems by creating smaller, manageable replicas that recreate critical characteristics of the original. In hydrology and civil infrastructure, for example, scaled models of dams help analyze water flow, turbulence, and pressure. This method allows for accurate predictions of real-world behavior within a controlled environment, significantly reducing the cost and time involved in full-scale...
249
The Availability Heuristic
5.9K
A heuristic is a general problem-solving framework (Tversky & Kahneman, 1974). You can think of these as mental shortcuts that are used to solve problems. Different types of heuristics are used in different types of situations, and the impulse to use a heuristic occurs when one of five conditions is met (Pratkanis, 1989):
5.9K
The Representativeness Heuristic
15.8K
The representative heuristic describes a biased way of thinking, in which you unintentionally stereotype someone or something. For example, you may assume that your professors spend their free time reading books and engaging in intellectual conversation, because the idea of them spending their time playing volleyball or visiting an amusement park does not fit in with your stereotypes of professors.
15.8K
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Typical Model Studies
344
Fluid mechanics model studies often utilize scaled-down systems to predict fluid behavior in full-scale environments, such as river flows, dam spillways, and structures interacting with open surfaces. Maintaining Froude number similarity in river models is crucial, as it replicates surface flow features like wave patterns and velocities.
344


