聊天GPT和参考间隔:对GPT-3.5 Turbo,GPT-4和GPT-4o中可重复性的比较分析
Annika Meyer1,2, Edgar Schömig3, Thomas Streichert2
1Department of Anesthesiology and Operative Intensive Care, Faculty of Medicine and University Hospital, University Hospital Cologne, Cologne, Germany.
像ChatGPT这样的大型语言模型在实验室医学中表现有前途,但在一致的参考间隔上扎. 新版本有所改进,但变化仍然存在,特别是在非标准化测试中.
科学领域:
- 医疗保健中的人工智能
- 实验室医学和诊断 实验室医学和诊断
- 临床病理学和信息学
背景情况:
- 大型语言模型 (LLM) 为实验室医学中的快速临床咨询提供了潜力.
- 关于LLM生成的参考间隔的一致性和临床可靠性存在不确定性,特别是没有临床背景.
研究的目的:
- 评估来自三个ChatGPT版本 (GPT-3.5-Turbo,GPT-4,GPT-4o) 的参考区间输出的可重复性.
- 通过使用参考区间变化作为应力测试来评估模型的一致性,当提示时省略区间信息.
主要方法:
- 一项涉及72万6000个聊天机器人请求与标准化提示的横截面研究.
- 分析了47个实验室参数中的246,842个参考间隔,以求一致性.
- 使用变量系数 (CV) 和回归模型来评估变量的统计分析.
主要成果:
- 参考区间的平均CV为26.50% (下限) 和15.82% (上限).
- 在GPT-4和GPT-4o中,CVs明显低于GPT-3.5-Turbo.
- 不一致的输出值得注意的是标准化不良的参数和不同的单位表达式.
结论:
- 虽然最新的ChatGPT版本显示了改进的重复性,但诊断上仍然存在不可接受的变异性,特别是在非标准化分析物中.
- 经过深思熟虑的快速设计,实验室实践的全球标准化,模型改进和监管监督至关重要.
- 当前的人工智能聊天机器人应该仅限于专业使用,并接受培训,在没有提供参考间隔的情况下拒绝解释.
更多相关视频
08:30Intraperitoneal Glucose Tolerance Test, Measurement of Lung Function, and Fixation of the Lung to Study the Impact of Obesity and Impaired Metabolism on Pulmonary Outcomes
Published on: March 15, 2018
09:30Pre-Implantation Genetic Testing for Aneuploidy on a Semiconductor Based Next-Generation Sequencing Platform
Published on: August 17, 2022
相关概念视频
Bioequivalence Data: Statistical Interpretation
Quantifying and Rejecting Outliers: The Grubbs Test
Comparing Experimental Results: Student's t-Test
Improving Translational Accuracy
Improving Translational Accuracy
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
