跨越3个临床专业的治疗建议的大型语言模型:比较研究
Theresa Isabelle Wilhelm1,2, Jonas Roos3, Robert Kaczmarczyk4,5
1Eye Center, Medical Center, Faculty of Medicine, University of Freiburg, Freiburg, Germany.
大型语言模型 (LLM) 在生成医疗信息方面表现有前途,但需要仔细评估准确性和安全性. 克劳德即时v1.0表现最好,而GPT-3.5-Turbo是最安全的,突出了医疗保健中持续AI评估的需要.
科学领域:
- 人工智能在医学中的应用
- 医疗信息学 医疗信息学
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 越来越多地用于医学信息生成.
- 严格评估人工智能产生的医疗内容的质量,准确性和安全性是跨专业的关键.
- 人工智能的快速发展需要评估其在医疗保健中的作用.
研究的目的:
- 为了评估四个著名的LLM的医疗内容生成性能:Claude-instant-v1.0,GPT-3.5-Turbo,Command-xlarge-nightly和Bloomz.
- 评估人工智能产生的眼科,骨科和皮肤病学治疗建议.
- 将医生评估与LLM生成的医学内容的自动GPT-4评估进行比较.
主要方法:
- 医生对使用mDISCERN评分,正确性和有害性的60种疾病的AI生成的治疗建议进行评估.
- 统计分析 (ANOVA,t测试) 用于比较模型和专业性能.
- 使用GPT-4的自动评估,与通过Pearson相关性对医生评估进行比较.
主要成果:
- 克劳德-即时-v1.0获得了最高的mDISCERN分数;布卢姆斯获得了最低的分数.
- 在LLM和专业之间观察到内容质量和安全性的显著差异.
- GPT-3.5-Turbo显示了最低的有害性评级;常见的错误包括诊断混乱和遗漏的治疗.
- GPT-4评估显示与医生评估有很大的一致性.
结论:
- 法律学证明了在产生医疗内容方面的能力,但质量和安全性需要改进.
- 定期,有方法的评估和监督对于可靠的AI医疗建议至关重要.
- GPT-4提供了一种可扩展的方法,用于对人工智能生成的医疗内容进行自动化,域异的评估.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
相关概念视频
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Language and Cognition
Treatment Strategies for Psychological Disorders
Psychological therapies focus on modifying emotions, thoughts, and behaviors through talking, interpreting, listening, rewarding, challenging, and modeling. Clinical psychologists, counselors, and social workers commonly practice psychotherapy. Clinical...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Typical Model Studies
Components of Language
