大型语言模型在临床预测任务中效率低于本地训练的机器学习模型
Katherine E Brown1, Chao Yan1, Zhuohang Li2
1Department of Biomedical Informatics, Vanderbilt University Medical Center (VUMC), Nashville, TN 37203, United States.
概括
传统的机器学习 (ML) 在使用电子健康记录 (EHR) 的临床预测任务中明显优于GPT-3.5和GPT-4等大型语言模型 (LLM). 与传统的ML方法相比,LLM在性能,校准和隐私保护方面存在局限性.
科学领域:
- 医疗保健中的人工智能
- 临床预测建模临床预测建模
- 电子健康记录 (EHR) 数据分析 数据分析
背景情况:
- 大型语言模型 (LLM) 越来越多地被用于各种应用,包括医疗保健.
- 传统的机器学习 (ML) 方法是使用电子健康记录 (EHR) 来进行临床预测的.
- 评估LLM作为临床环境中传统ML的潜在替代品,对于推进医疗保健技术至关重要.
研究的目的:
- 与传统的ML相比,评估当前LLM (GPT-3.5,GPT-4) 作为临床预测者的有效性.
- 调查影响临床预测中LLM采用的因素:性能,校准,公平性和隐私弹性.
- 分析数据概括对隐私对模型性能的影响.
主要方法:
- 在VUMC和MIMIC IV.的EHR数据上对GPT-3.5,GPT-4和梯度增强树 (传统ML) 的比较分析.
- 通过接收器操作特征下的面积 (AUROC) 和校准的障碍得分来测量性能.
- 公平性评估使用均等赔率和跨人口群体的统计平价.
- 评估了在AUROC上环境学习和数据概括对隐私的影响.
主要成果:
- 与GPT-3.5 (AUROC:0.537,0.517) 和GPT-4 (AUROC:0.629,0.602) 相比,传统的ML表现出明显优异的预测性能 (AUROC:0.847,0.894). 这两种方法都具有较高的预测性能.
- 传统的ML还显示出比GPT-3.5 (0.384,0.06) 和GPT-4 (0.251,0.219) 更好的输出概率校准 (障碍得分:0.134,0.042).
- GPT-4显示,在人口群体之间,公平性更好,但以模型性能降低为代价.
结论:
- 目前,非微调的LLM在使用EHR数据进行临床预测方面,效果和稳定性比本地训练的ML要低.
- 传统的ML模型在将人口统计信息用于隐私保护方面更强大.
- 虽然目前的LLMs落后于传统的ML,但持续的进步表明未来临床应用的潜力.
更多相关视频
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
8.2K
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.3K
相关概念视频
Language and Cognition
308
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
308
Sensitivity, Specificity, and Predicted Value
157
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
157
Improving Translational Accuracy
2.5K
2.5K
