用大型语言模型检测科学文献中的参考错误
Tianmai M Zhang1, Neil F Abernethy1
1Department of Biomedical Informatics and Medical Education, School of Medicine, University of Washington, Seattle, WA, USA.
大型语言模型可以检测科学论文中的引用错误,即使信息有限. 这种人工智能的进步有助于确保科学文献的完整性和准确的信息传播.
科学领域:
- 人工智能的人工智能
- 学术出版学术出版
- 科学完整性 科学完整性
背景情况:
- 引用错误,如引用和引文错误,在科学出版物中很普遍.
- 这些错误可以传播错误信息,并且难以手动识别,威胁到科学文献的完整性.
- 需要自动检测方法来应对这些挑战.
研究的目的:
- 评估大型语言模型 (LLM) 在科学文章中检测引文错误方面的有效性.
- 通过检索增强,通过不同级别的上下文信息来评估LLM绩效.
主要方法:
- 开发一个由专家注释的数据集,包括来自期刊文章的陈述-参考对,具有重要的生物医学组成部分.
- 在此数据集上对OpenAI的大型语言模型GPT家族的评估.
- 在各种环境中测试LLM,包括那些参考数据有限的环境.
主要成果:
- 大型语言模型表现出了显著的识别错误引用的能力.
- 即使在有限的上下文信息和没有模型微调的情况下,也可以实现有效的检测.
- 这项研究证实了人工智能在支持科学写作和审查过程中的潜力.
结论:
- 大型语言模型有望自动检测科学文献中的引用错误.
- 人工智能工具可以帮助保持发表的研究的准确性和可靠性.
- 这项研究有助于利用人工智能增强科学沟通并确保事实依据.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
相关概念视频
Improving Translational Accuracy
Improving Translational Accuracy
Detection of Gross Error: The Q Test
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
NMR Spectrometers: Resolution and Error Correction
