专家评估临床对话总结大型语言模型的专家评估
David Fraile Navarro1, Enrico Coiera2, Thomas W Hambly3
1Centre for Health Informatics, Australian Institute of Health Innovation, Macquarie University, Level 6, 75 Talavera Road, North Ryde, Sydney, NSW, 2113, Australia. david.frailenavarro@mq.edu.au.
Scientific reports
|January 8, 2025
概括
像ChatGPT这样的大型语言模型在总结临床对话方面表现有希望,接近人类质量. 然而,ROUGE指标对于临床文本评估可能不可靠.
科学领域:
- 人工智能的人工智能
- 临床信息学 临床信息学
- 自然语言处理自然语言处理.
背景情况:
- 临床对话的自动总结对于有效的医疗保健文档至关重要.
- 在这个领域评估大型语言模型 (LLM) 的性能需要强大的指标和人类监督.
研究的目的:
- 评估和比较各种LLMs在总结临床对话中的表现.
- 评估计算指标 (ROUGE,UniEval) 与专家人类评估的可靠性,以生成临床总结.
主要方法:
- 五个LLM的探索性评估,包括一般和微调模型,以及ChatGPT.
- 使用ROUGE和UniEval指标进行评估.
- 专家临床医生评估,将模型生成的摘要与人类生成的黄金标准进行比较.
主要成果:
- 在UniEval和临床医师评估中,ChatGPT在连贯性,一致性,流性和整体临床实用性方面获得了最高分.
- 一个微调的变压器模型在ROUGE指标上表现最好,而ChatGPT表现最低.
- 与ROUGE不同,UniEval与人类评级有很强的相关性,这表明它对临床摘要的可靠性更高.
结论:
- 在总结临床对话方面,ChatGPT的表现接近人类质量.
- 对于评估自动化临床摘要生成,UniEval是一个比ROUGE更可靠的指标.
- 法律法规提供了一个潜在的解决方案,用于自动化临床对话总结,但隐私和数据访问仍然是重大挑战.
更多相关视频
相关概念视频
Improving Translational Accuracy
2.5K
2.5K
Techniques of Therapeutic Communication II: Focusing, Paraphrasing, and Summarizing
7.7K
Focusing involves centering a conversation on a message's critical elements or concepts. Focusing is valuable if the talk is vague or patients begin to repeat themselves. Sometimes, when patients are asked about their symptoms, they may go off-topic and try to tell their entire life story. Respectfully, the nurse should bring the conversation back into focus.
This therapeutic technique can also be used when a patient brings up pertinent information during a health-related conversation. The...
This therapeutic technique can also be used when a patient brings up pertinent information during a health-related conversation. The...
7.7K


