在电子健康记录应用程序中测试和评估生成型大语言模型:系统审查
Xinsong Du1,2,3, Zhengyang Zhou4, Yifei Wang4
1Division of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, MA 02115, United States.
生成型大语言模型 (LLM) 在医疗保健方面表现有前途,但需要在各个专业和任务中进行更广泛的评估. 目前的指标与人体表现没有很强的相关性,因此需要标准化的框架来有效的临床应用.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 临床数据分析 临床数据分析
背景情况:
- 生成型大语言模型 (LLM) 越来越多地与电子健康记录 (EHR) 数据一起用于临床和研究支持.
- 本系统性审查旨在描述当前在医疗保健中LLM的应用和评估.
研究的目的:
- 系统地审查和描述临床领域和使用用EHR数据分析的生成LLM的用例.
- 总结评估方法,并确定这些研究中使用的常见的LLM和指标.
主要方法:
- 按照PRISMA指南进行系统审查,搜索PubMed和科学网络 (2023年1月 - 2024年11月).
- 纳入标准:生成的LLM分析现实世界EHR数据,并进行定量性绩效评估.
- 数据提取集中在临床专业,任务,使用的LLM和评估指标上.
主要成果:
- 196项研究符合标准,主要是放射学 (26.0%),瘤学 (10.7%) 和急诊医学 (6.6%).
- 临床决策支持是最常见的任务 (62.2%);总结和患者沟通是最不常见的.
- GPT-4和GPT-3.5是主导的LLM;确定了22个非NLP和35个NLP指标,NLP指标与人类评估没有很强的相关性.
结论:
- 需要对跨不同临床专业和任务的生成性LLM进行更广泛的评估.
- 迫切需要在EHR数据中开发标准化,可扩展和临床上有意义的LLM评估框架.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
相关概念视频
Methods of Documentation VII: EMR
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Nursing Evaluation
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
