评估临床AI总结与大型语言模型作为法官
Emma Croxford1, Yanjun Gao2, Elliot First3
1Department of Biostatistics and Medical Informatics, University of Wisconsin, Madison, USA.
NPJ digital medicine
|November 5, 2025
概括
使用大型语言模型 (LLM) 方法自动评估人工智能生成的临床摘要,可显著提高准确性评估. 这种方法为电子健康记录 (EHR) 的手动审查提供了一个可扩展和高效的替代方案.
科学领域:
- 人工智能在医学中的应用
- 临床信息学 临床信息学
- 自然语言处理自然语言处理.
背景情况:
- 电子健康记录 (EHR) 产生了大量的临床数据,这给医疗保健提供者带来了高效合成的挑战.
- 生成型人工智能和大型语言模型 (LLM) 提供了概括EHR的潜力,以减少提供者的认知负载.
- 确保人工智能生成的摘要的准确性需要强大的评估方法,因为人类审查是耗时和昂贵的.
研究的目的:
- 引入和验证基于LLM的自动化方法,以评估从现实世界EHR数据中获得的多文档摘要的质量.
- 将LLM-as-a-Judge框架与已建立的提供者文档总结质量工具 (PDSQI) 进行比较.
主要方法:
- 开发和验证了一个LLM-as-a-Judge框架,用于评估EHR多文档摘要.
- 基于LLM的评估与使用PDSQI的人类审查进行了基准评估.
- 评估评价者之间的可靠性和绩效指标,包括类内相关系数 (ICC) 和评估时间.
主要成果:
- 作为法官的法学士框架表现出强大的评审者之间的可靠性,与人类评审者相美.
- 从人类评估中,GPT-o3-mini获得了0.818的ICC和0的中位分差,在22秒内完成了评估.
- 推理的LLM在评级者之间的可靠性方面表现优于其他方法,特别是在需要领域专业知识的复杂评估中.
结论:
- 基于LLM的自动化评估提供了一个可扩展和高效的方法来评估AI生成的临床摘要的准确性和安全性.
- 作为法官的法学士 (LLM-as-a-Judge) 方法可以显著降低与传统的人类审查相关的负担和成本.
- 这项技术有助于可靠地部署人工智能来合成电子健康记录中的临床数据.
相关概念视频
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Clinical Trials: Overview
4.5K
Clinical development focuses on how the drug will interact with the human body and encompasses four key phases of clinical trials, each serving a specific purpose in assessing the safety and effectiveness of new drugs. These phases overlap and build upon one another. Phase I involves a small group of healthy volunteers (typically 20-80 individuals) or, in cases where significant toxicity is expected, patients with the targeted disease, such as cancer or AIDS. The volunteers are tested for...
4.5K

