ECAsT:用于对话搜索的大数据集和衡量度量强度的评估
Haya Al-Thani1, Bernard J Jansen2, Tamer Elsayed3
1College of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar.
PeerJ. Computer science
|June 22, 2023
概括
研究人员创建了Expanded-CAsT (ECAsT),这是一个比以前大665%的对话搜索数据集. 这一数据集增强了神经模型训练,并减少了对话性搜索系统中的评估偏差.
科学领域:
- 信息检索 信息检索
- 自然语言处理自然语言处理.
- 人工智能的人工智能
背景情况:
- 文本检索会议对话辅助轨道 (CAST) 基准对话段落检索,但使用有限的数据集.
- 现有的CAST数据集包含1000多个转和100个对话主题,阻碍了大规模的基准测试.
研究的目的:
- 通过创建一个更大,更多样化的数据集来解决对话搜索基准测试中的数据集限制.
- 评估使用新数据集的传统对话评估指标的稳定性和识别偏差.
主要方法:
- 开发了Expanded-CAsT (ECAsT),一种新的多轮对话数据集,使用多阶段查询重构和神经转述.
- 创建了一个新的模式,用于生成多转句,通过人类和自动评估验证意义和多样性.
- 通过ECAST数据集,评估传统的CAST评估指标的稳定性和对语言多样性的偏见.
主要成果:
- 创建了ECAST,一个对话式搜索数据集,拥有9200多个轮回,代表了语言规模和多样性的665%的增加.
- 证明通过表述语法将语言多样性纳入其中,可以提高24%的段落检索,而CAST基线是2%.
- 鉴定了传统指标中的偏见,并显示了语言多样性的好处,减少了评估偏见和改善了聚合段落收集.
结论:
- 扩展的ECAsT数据集显著增强了对话式搜索模型培训和测试,由于增加了规模和语言多样性.
- 该研究强调了对话评估中传统指标的局限性,并强调了语言多样性对公正评估的重要性.
- ECAsT为研究界提供了一种宝贵的资源,以推进大规模的开放域对话搜索基准测试.
更多相关视频
05:48Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.5K
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
332
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Data Collection by Experiments
24.4K
Data collection is a systematic method of obtaining, observing, measuring, and analyzing accurate information. An experimental study is a standard method of data collection that involves the manipulation of the samples by applying some form of treatment prior to data collection. It refers to manipulating one variable to determine its changes on another variable. The sample subjected to treatment is known as “experimental units.”
An example of the experimental method is a public...
An example of the experimental method is a public...
24.4K
Data Collection by Observations
12.1K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
12.1K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Data Collection by Survey
6.5K
The systematic method of obtaining and analyzing accurate information of a population is called data collection. A survey is a standard method of data collection that involves collecting information from a target human population about their experience, opinion, or knowledge of a product, service, or process. The responses are recorded and interpreted. The most common survey examples are written questionnaires, face-to-face or telephonic conversations, focus groups, and electronic (e-mail or...
6.5K
Confidence Coefficient
7.7K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
7.7K
