比较大型语言模型的准确性和即时工程在诊断现实世界的案例
Guanhong Yao1, WuJi Zhang2, Yingxi Zhu2
1School of Information Engineering, Beijing Polytechnic College, Beijing, China.
大型语言模型 (LLM) 显示出从医疗记录诊断疾病的前景. 短镜头提示提高了准确性,但增加了成本,突出了优化需求.
科学领域:
- 人工智能在医学中的应用
- 临床决策支持系统 临床决策支持系统
- 在医疗保健中的自然语言处理.
背景情况:
- 大型语言模型 (LLM) 为临床决策提供了潜力,特别是在诊断复杂和罕见疾病方面.
- 在医疗保健中LLM的现实应用需要对其诊断准确性和实用性的严格评估.
研究的目的:
- 为了评估四种LLM的诊断性能:GPT-4o mini,GPT-4o,ERNIE和Llama-3.
- 通过使用现实世界的住院病人的医疗记录,评估一些射击和思维链提示对LLM诊断准确性的影响.
主要方法:
- 北京大学国际医院的一项回顾性研究分析了1,122份医疗记录.
- 记录包括常见/罕见的风湿性自身免疫性疾病和非风湿性疾病.
- 诊断准确性 (hit1) 通过将正确的诊断纳入顶级预测来衡量,并使用少数镜头和思维链提示进行评估.
主要成果:
- 基本的LLM hit1率从81.8%到82.9%不等.
- 少数射击提示显著改善了GPT-4o的hit1到85.9% (p=0.02),优于其他模型 (p<0.05).
- 连锁思维提示没有产生显著的改善;对于类风湿性疾病的诊断准确性高于非类风湿性疾病. 少数射击提示增加了每次诊断的GPT-4o成本约4.54.0日元.
结论:
- 包括GPT-4o在内的LLM在真实世界医学数据上表现出相当大的诊断准确性.
- 低射击提示提高了LLM的性能,但造成了更高的成本,需要成本效益分析和进一步提高准确性.
- 研究结果支持在中国医疗环境中LLM的发展,并强调需要多中心验证.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
相关概念视频
Improving Translational Accuracy
Language and Cognition
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Detection of Gross Error: The Q Test
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
