评估大型语言模型,用于起草紧急部门遭遇总结
Christopher Y K Williams1, Jaskaran Bains2, Tianyu Tang2
1Bakar Computational Health Sciences Institute, University of California, San Francisco, California, United States of America.
大型语言模型 (LLM) 可以总结临床笔记,但会出现幻觉和遗漏等错误. 虽然大多数错误的危害潜力很低,但对患者安全而言,临床医生的仔细审查是必不可少的.
科学领域:
- 临床信息学 临床信息学
- 医疗保健中的人工智能
- 医疗文件 医疗文件
背景情况:
- 大型语言模型 (LLM) 对文本总结等临床应用具有前景.
- 越来越多的人工智能写字器的部署需要在医疗保健环境中严格评估它们的准确性.
研究的目的:
- 评估GPT-4和GPT-3.5-turbo在生成紧急部门 (ED) 遭遇总结中的性能.
- 在LLM生成的ED摘要中识别错误的普遍性和类型 (不准确性,幻觉,遗漏).
主要方法:
- 对100个随机抽样成年ED访问 (2012-2023) 的横截面研究.
- 对GPT-4和GPT-3.5-turbo的评估根据三个标准生成了摘要:不准确性,幻觉和遗漏.
- 在遭遇总结部分中分析错误类型和位置.
主要成果:
- 在33%的案例中,GPT-4生成了无错误的摘要;在10%的案例中,GPT-3.5-turbo生成了无错误的摘要.
- 在GPT-4总结中,不准确率为10%,幻觉率为42%,遗漏率为47%.
- 在"计划"部分,不准确/幻觉是常见的;在"体检"和"投诉提交历史"部分,遗漏是常见的.
- 对错误的平均潜在危害得分很低 (0.57/7).
结论:
- LLM可以生成临床经验总结,但容易产生幻觉和遗漏.
- 在LLM生成的临床文本中的错误,虽然往往是低危害的,但需要临床医生仔细审查.
- 了解错误模式对于安全地将LLMs集成到临床工作流程中至关重要.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
相关概念视频
Methods of Documentation VII: EMR
Types of Reports III: Telephone and Verbal Reports
Here's an overview of each type:
Telephone Orders
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
SBAR II: Application of SBAR
SBAR Report from a Nurse to a Health Care Provider
S: "Hello, Dr. Smith. This is Jane, RN, from the Med Surg unit. I am calling to tell you about Ms. White in Room 210, who is experiencing increased pain and redness at her incision site. Her recent...
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
