评估ChatGPT模型版本的准确性,以提供寻求护理的建议
Marvin Kopka1, Longqi He2, Markus A Feufel2
1Division of Ergonomics, Department of Psychology and Ergonomics (IPA), Technische Universität Berlin, Berlin, Germany. marvin.kopka@tu-berlin.de.
聊天GPT模型在寻求护理建议方面显示出有限的准确性. 虽然新版本提供了自我护理建议,但它们的可靠性不足以独立使用. 聚合输出可以提高绩效.
科学领域:
- 医疗保健中的人工智能
- 医疗决策支持系统 医疗决策支持系统
- 自然语言处理自然语言处理.
背景情况:
- 非专业人士越来越多地使用像ChatGPT这样的AI工具来做健康决策.
- 人工智能产生的寻求护理建议的准确性,特别是来自较新的模型,需要进行彻底的评估.
- 现有的研究还没有全面评估所有可用的ChatGPT模型的医疗建议准确性.
研究的目的:
- 评估目前所有可用的ChatGPT模型提供的寻求护理建议的准确性.
- 为了比较不同版本的ChatGPT在医疗紧急情况分类中的性能.
- 探索改善人工智能产生的医疗建议准确性的方法.
主要方法:
- 评估了22个ChatGPT模型,使用了45个验证的标签,每个标签被提示了10次 (共9900次评估).
- 紧急情况分为紧急护理,非紧急护理或自我护理类别.
- 根据两名医生建立的黄金标准验证了AI建议,并测试了聚合算法.
主要成果:
- o1-迷你模型以74%的精度显示出最高的精度.
- 在较新的ChatGPT模型中没有观察到一致的准确性改善.
- 推理模型显示自我护理案例识别得到了改进;汇总输出的准确性提高了4%.
- 在多个试验中选择最低的紧急级别提高了整体准确性.
结论:
- 目前的ChatGPT模型在寻求护理的决策中无法独立使用,因为不够准确.
- 新的AI模型越来越多地提供自我护理建议,但可靠性仍然是一个问题.
- 使用聚合算法来利用输出可变性可以显著提高现有的AI模型的性能.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
相关概念视频
Accuracy, limits, and approximation
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Improving Translational Accuracy
Improving Translational Accuracy
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Accuracy and Precision
