人与人工智能:比较Cochrane作者和ChatGPT的偏差评估风险
1Cochrane Evidence Synthesis Unit Germany/UK, Institut für Allgemeinmedizin (ifam), Universitätsklinikum Düsseldorf Heinrich-Heine-Universität Düsseldorf Germany.
Cochrane evidence synthesis and methods
|September 5, 2025
概括
在随机对照试验 (RCT) 中评估偏差风险时,ChatGPT-4o与人类评审者之间存在中度的共识. 虽然人工智能在证据综合方面表现有前途,但在系统审查中需要进一步改进以获得最佳表现.
科学领域:
- 医疗信息学
- 医疗保健中的人工智能
- 证据综合
背景情况:
- 系统性审查和元分析对于临床决策至关重要,但需要大量的时间.
- 人工智能 (AI) 提供了加速证据合成过程的潜力.
- 偏差风险评估是系统审查的关键组成部分.
研究的目的:
- 使用偏差风险2 (RoB2) 工具评估ChatGPT-4o对随机对照试验 (RCT) 的偏差风险的性能.
- 将人工智能驱动的偏见风险评估与科克兰评论中人类审查员进行的风险评估进行比较.
- 在RoB 2评估中量化ChatGPT-4o的一致性和准确性.
主要方法:
- 采用RoB 2工具选择了可克兰评价的样本.
- ChatGPT-4o被要求评估基于RoB 2域的RCT偏差风险.
- 使用加权的卡帕统计数据,以及准确性,敏感性和特异性来衡量一致性.
主要成果:
- 对于偏差判断的整体风险,ChatGPT-4o与人类审查者达成了中等一致 (加权kappa=0. 51).
- 根据领域的不同,协议从公平 (报告结果的选择) 到适度 (结果的测量) 不同.
- 在高风险研究中,人工智能显示了53%的敏感性,在低风险研究中显示了99%的特异性.
结论:
- 在使用RoB2工具进行偏差风险评估时,ChatGPT-4o表现出相当或中等的能力.
- 人工智能辅助偏差风险评估显示出潜力,但需要进一步开发和快速工程.
- 未来的研究应该专注于标准化提示和相互评价的可靠性,以便进行更强大的比较.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
Systematic Error: Methodological and Sampling Errors
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
Accuracy and Errors in Hypothesis Testing
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...


