利用人工智能进行元分析:评估LLM在检测下一代证据综合的出版偏见
Xing Xing1, Lifeng Lin2, Mohammad Hassan Murad3
1Department of Biostatistics Johns Hopkins Bloomberg School of Public Health Baltimore Maryland USA.
Cochrane evidence synthesis and methods
|September 29, 2025
概括
大型语言模型 (LLM) 显示在元分析中检测出版偏差 (PB) 的能力有限. 在LLM能够可靠地支持系统性审查和证据综合之前,需要进行专门的调整.
科学领域:
- 生物医学信息学 生物医学信息学
- 数据科学数据科学数据科学
- 进行元分析的方法论.
背景情况:
- 出版偏差 (PB) 通过歪曲效应大小估计,损害了元分析的有效性.
- 大型语言模型 (LLM) 具有先进的模式识别和多式联络能力.
- 法律法规可能会提高系统审查效率和PB评估.
研究的目的:
- 评估最先进的多式联络LLM在检测出版偏差方面的有效性.
- 评估GPT-4o和Llama 3.2视觉在使用漏斗图和定量数据识别PB方面的表现.
主要方法:
- 在不同条件下的模拟元分析:没有PB,不同的PB水平,研究数量和异质性.
- 仅使用漏斗图表和使用定量输入来评估LLM绩效.
- 比较了GPT-4o和Llama 3.2视觉检测出版偏差的能力.
主要成果:
- 无论是GPT-4o还是Llama 3.2视觉都没有在所有测试场景中始终检测到出版偏差.
- 在没有PB条件下,GPT-4o表现出比Llama 3.2 Vision更高的特异性,特别是在>=20项研究中.
- 量化输入,异质性和未报告的研究模式并没有显著改善LLM的表现.
结论:
- 目前的LLM缺乏可靠地检测出版偏差的能力,而不需要微调.
- 专门的模型适应对于将LLMs集成到元分析工作流程中至关重要.
- 未来的研究应该专注于有针对性的LLM改进,以改善证据合成.
相关概念视频
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K

