大型语言模型在一些系统审查任务中显示出有前途的表现,但需要谨慎实施:系统审查
Florian Laignelot1, Guillaume L Martin1, Mohamad Ossman1
1Sorbonne Université, INSERM, Institut Pierre Louis d'Epidémiologie et de Santé Publique, UMR-S 1136, AP-HP, Hôpital Pitié-Salpêtrière, Département de Santé Publique, Paris, France.
Journal of clinical epidemiology
|March 14, 2026
概括
大型语言模型 (LLM) 显示出对自动化系统审查任务 (如选) 的承诺,新型模型的性能更好. 仔细实施是将LLMs纳入系统审查的关键.
科学领域:
- 生物医学信息学是生物医学信息学.
- 医疗保健中的人工智能
- 系统审查方法论系统审查方法论
背景情况:
- 越来越多的生物医学文献给进行系统审查带来了挑战.
- 需要自动化系统审查流程以提高效率.
研究的目的:
- 评估大型语言模型 (LLM) 在自动化系统审查和元分析步骤中的性能.
- 评估系统审查工作流程的各个阶段的LLM能力.
主要方法:
- 在系统审查中对评估LLM绩效的研究进行系统审查.
- 在2025年1月14日之前搜索了PubMed,Embase,Cochrane图书馆和预印平台.
- 提取数据并评估偏差风险,分析积极 (PPA) 和负面百分比协议 (NPA) 等绩效指标.
主要成果:
- 包括63项研究与148个LLM绩效评估,主要是GPT模型.
- 在标题/摘要 (PPA 0.92,NPA 0.89) 和全文选 (PPA 0.93,NPA 0.92) 中,LLM表现出很高的表现.
- 数据提取 (中位准确率0.95) 和偏差风险评估 (中位准确率0.62) 显示出可变但有希望的结果,新型LLM的表现优于旧型LLM.
结论:
- 大型语言模型显示出在系统性审查,特别是选中,自动化重复任务的巨大潜力.
- 整合LLMs需要仔细实施和适当的保障,以确保在证据综合中可靠使用.
相关概念视频
Improving Translational Accuracy
15.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.4K
Improving Translational Accuracy
3.7K
3.7K


