TG-CSR:一个基于九个正式的常识类别的人类标记数据集
Henrique Santos1, Alice M Mulvehill1, Ke Shen2
1Rensselaer Polytechnic Institute 110 8th St., Troy, NY 12180, USA.
Data in brief
|October 25, 2023
概括
本研究详细介绍了理论基础的常识推理 (TG-CSR) 基准的人类注释标签. 释放这些数据可以分析机器常识推理和注释器一致性.
科学领域:
- 人工智能的人工智能
- 自然语言处理自然语言处理.
- 认知科学 认知科学
背景情况:
- 机器常识推理 (MCS) 旨在在机器中复制类似人类的决策.
- 基准和问答数据集对于评估MCS模型进展至关重要.
- 现有的基准标准通常在数据创建过程中缺乏透明度.
研究的目的:
- 描述TG-CSR基准值的人类注释器生成的个别标签数据.
- 为了促进对注释过程的分析,包括注释者之间的协议和噪音.
- 为推进机器常识推理研究提供透明的数据集.
主要方法:
- 六名人类注释者获得了TG-CSR提示和具体说明.
- 在结构化注释会话期间,标签被插入到电子表格单元中.
- 数据被组织成JSON,电子表格和JSONL文件格式.
主要成果:
- 从人类注释者收集了个人原始和规范化的标签数据.
- 数据集包括对TG-CSR的地面真相创建的详细记录.
- 数据结构支持详细检查标签过程.
结论:
- 发布个人注释标签可以提高基准指标开发的透明度.
- 这些数据使得对人工智能人为标签的噪音和一致性进行关键研究成为可能.
- 促进对注释过程的更深入分析将改善未来的常识推理基准.
更多相关视频
相关概念视频
Natural and Artificial Concepts
171
In psychology, concepts can be divided into two categories: natural and artificial. Natural concepts are formed through direct or indirect experiences. For example, consider the concept of snow. If you live in a place with regular snowfall, such as Essex Junction, Vermont, you know snow through direct experiences. You’ve seen it fall, touched it, shoveled it, and played in it. You recognize its texture, appearance, and even its smell. In contrast, if you live on an island like Saint...
171
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K
Deductive Reasoning
55.3K
Deductive reasoning, or deduction, is the type of logic used in hypothesis-based science. In deductive reasoning, the pattern of thinking moves in the opposite direction as compared to inductive reasoning, which means that it uses a general principle or law to predict specific results. From those general principles, a scientist can deduce and predict the specific results that would be valid as long as the general principles are valid.
For example, a researcher can deduce specific predictions...
For example, a researcher can deduce specific predictions...
55.3K
How Data are Classified: Categorical Data
33.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
33.8K
Inductive Reasoning
60.5K
Inductive reasoning is a form of logical thinking that uses related observations to arrive at a general conclusion. It is uncertain and operates in degrees to which the conclusions are credible. As such, inductive arguments can be weak or strong, rather than valid or invalid, and conclusions can be used to formulate testable, falsifiable hypotheses.
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
60.5K
Data: Types and Distribution
731
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
731


