在引用选中LLM的表现:与各种包含率的数据集进行比较
Zhihong Zhang1,2, Mohamad Javad Momeni Nezhad3, Pallavi Gupta2
1Data Science Institute, Columbia University, New York, NY 10027, United States.
Studies in health technology and informatics
|August 8, 2025
概括
大型语言模型 (LLM) 显示出在系统审查中自动化引用选的前景. 合并方法和多数投票改善了跨不同数据集的LLM性能.
科学领域:
- 生物医学信息学 生物医学信息学
- 研究中的人工智能.
- 证据综合 证据综合
背景情况:
- 系统性审查需要大量的手工工作来确定研究.
- 自动化引用选可以显著加快审查过程.
研究的目的:
- 评估大型语言模型 (LLM) 在自动引用选中的有效性.
- 评估跨不同研究纳入率的数据集的LLM绩效.
主要方法:
- 测试了六个LLM使用零到五次射击的上下文学习.
- 使用PubMedBERT进行基于语义相似性的示范选择.
- 利用多数投票和集体学习来提高选准确度.
主要成果:
- 没有一个LLM在所有数据集中始终表现优于其他LLM.
- 根据数据集包含率,LLM的敏感性和特异性各不相同.
- 集体学习和多数投票表明,引文选性能有所改善.
结论:
- 在系统性审查中,LLM提供了一个可行的工具,用于自动化引用选.
- 整体方法对于优化LLM在这个任务中的表现至关重要.
- 需要进一步的研究来完善在各种系统性审查场景中LLM的应用.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
579
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.6K
相关概念视频
Sample Size Calculation
3.8K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
3.8K
Improving Translational Accuracy
2.7K
2.7K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Goodness-of-Fit Test
4.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.1K
Quantifying and Rejecting Outliers: The Grubbs Test
2.0K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.0K
Statistical Significance
20.4K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
20.4K
