在系统性审查中,ChatGPT-4o可以作为数据提取的第二评分器
Mette Motzfeldt Jensen1,2, Mathias Brix Danielsen1,2, Johannes Riis1,2
1Department of Geriatric Medicine, Aalborg University Hospital, Aalborg, Denmark.
PloS one
|January 8, 2025
概括
像ChatGPT-4o这样的人工智能 (AI) 在系统审查中显示出数据提取的高有效性和可重复性. 这种人工智能工具可以作为第二个审查员,简化临床指南的证据综合.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 基于证据的医学基于证据的医学.
背景情况:
- 系统审查对于证据综合至关重要,但需要大量的时间.
- 人工智能 (AI) 工具,如ChatGPT-4o,有可能加速数据提取过程.
- 在系统性审查中验证人工智能的有效性至关重要.
研究的目的:
- 评估ChatGPT-4o在数据提取方面的有效性.
- 为了评估ChatGPT-4o的数据提取能力的可复制性.
主要方法:
- 通过ChatGPT-4o和两个独立的人类审查员从系统审查论文中提取的数据的比较分析.
- 使用五类规模 (完全正确到错误数据) 的有效性评估.
- 通过使用不同的ChatGPT-4o帐户进行单独的数据提取会话进行可重复性测试.
主要成果:
- 在11篇论文中,ChatGPT-4o在数据提取方面实现了92.4%的准确性,其中5.2%的数据是错误的.
- 在提取过程之间观察到高的整体可重复性 (94.1%).
- 对于在论文中未明确报道的数据,可复制性降至77.2%.
结论:
- 在系统审查中,ChatGPT-4o在数据提取方面表现出高的有效性和可重复性.
- 人工智能工具适合在系统审查中作为第二审核员使用.
- 聊天GPT-4o显示了未来在证据合成数据总结方面的进展的希望.
相关概念视频
Detection of Gross Error: The Q Test
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Quantifying and Rejecting Outliers: The Grubbs Test
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...


