评估ChatGPT-5,Claude和DeepSeek在医学统计分析中的性能
1University Institute for Primary Care (IuMFE), University of Geneva, 1, Rue Michel-Servet, 1211, Geneva 4, Switzerland. paulsebo@hotmail.com.
Internal and emergency medicine
|February 26, 2026
概括
大型语言模型 (LLM) 在识别统计测试和生成用于生物医学研究的Stata命令方面表现出高准确性. 这些先进的AI工具提供可靠的统计支持,错误风险最小.
科学领域:
- 生物医学研究的研究.
- 人工智能的人工智能是人工智能.
- 统计分析 统计分析
背景情况:
- 大型语言模型 (LLM) 在生物医学研究中越来越多地用于统计支持.
- 在选择统计测试和生成软件命令时,LLM的可靠性需要彻底评估.
研究的目的:
- 为了比较ChatGPT-5,Claude和DeepSeek在统计测试选择和Stata命令生成中的性能.
- 评估LLM生成的统计命令的准确性和可重复性.
主要方法:
- 三十二个Stata教程示例被用来测试ChatGPT-5,Claude和DeepSeek.
- 响应被根据对参考命令 (COR,SYN,ALT,CMM) 的偏差进行分类.
- 准确性计算为没有或低风险偏差的输出比例.
主要成果:
- 所有模型在识别统计测试时都实现了100%的准确性.
- 在所有模型中观察到Stata命令生成的高精度 (90.6%-96.9%).
- 测试轮之间可重复性很好,高风险偏差率很低.
结论:
- 聊天GPT-5,Claude和DeepSeek在统计推理任务中表现出高准确度和可重复性.
- 罕见的高风险偏差表明,LLM可以成为应用统计推理的补充工具.
- 先进的LLM在支持生物医学研究中的统计分析方面表现有前途.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
661
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
661
Statistical Software for Data Analysis and Clinical Trials
1.7K
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.7K


