在证据合成中,用于数据提取的两个大型语言模型的性能
Amanda Konet1, Ian Thomas1, Gerald Gartlehner1,2
1Social, Statistical, and Environmental Sciences, RTI International, Research Triangle Park, North Carolina, USA.
Research synthesis methods
|June 19, 2024
概括
大型语言模型 (LLM) 显示了证据综合数据提取的前景. 克劳德2显示出比GPT-4更高的精度,这主要是由于GPT-4的优势.
科学领域:
- 科学研究中的人工智能
- 生物医学信息学 生物医学信息学
- 证据综合的方法论.
背景情况:
- 准确的数据提取对于可靠的证据合成至关重要.
- 大型语言模型 (LLM) 具有自动化数据提取的潜力.
- 关于证据综合任务的最佳LLM存在不确定性.
研究的目的:
- 为了比较两种广泛使用的LLM,Claude 2和GPT-4的性能,用于证据合成中的数据提取.
- 评估从全文科学文章中以LLM驱动的数据提取的准确性.
主要方法:
- 两种LLM (第2条,GPT-4) 用于从10篇已发表的文章中提取预先规定的数据元素.
- 使用LLMs的浏览器版本处理了完整的研究PDF.
- GPT-4需要第三方插件进行PDF解析,而Claude 2则没有.
- 通过将LLM输出与之前提取的数据进行比较来评估准确性.
主要成果:
- 克劳德2在数据提取方面实现了高精度 (96.3%).
- GPT-4及其PDF解析插件的准确性较低 (68.8%),大多数错误归因于插件.
- 两种LLM都准确地识别了缺失的数据元素,并处理了未报告的信息.
- 当提供精选文本时,Claude 2和GPT-4分别实现了98.7%和100%的准确性.
结论:
- 在证据合成中,LLM具有显著的潜力,可以提高数据提取效率.
- 准确的PDF解析是影响LLM性能的一个关键因素.
- 人类监督对于验证LLM生成的提取值仍然至关重要.
更多相关视频
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
440
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.2K
相关概念视频
Extraction: Advanced Methods
446
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
446
Improving Translational Accuracy
2.6K
2.6K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Crossover Experiments
2.8K
Crossover experiments, also called the repeated-measurements design, is a study design in which all experimental units are exposed to all treatments in different periods. Crossover experiments are generally used in psychology, the pharmaceutical industry, agriculture, and medicine.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
2.8K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
