一项概念验证研究,用于患者使用大型语言模型的开放笔记
Liz Salmi1,2, Dana M Lewis3, Jennifer L Clarke4
1Department of Women's and Children's Health, Uppsala University, 752 37 Uppsala, Sweden.
JAMIA open
|April 10, 2025
概括
大型语言模型 (LLM) 可以帮助患者理解临床笔记. 使用基于人格的提示与像ChatGPT 4o这样的LLM提高了患者查询响应的准确性和相关性.
科学领域:
- 医疗保健中的人工智能
- 临床信息学 临床信息学
- 患者参与技术的技术
背景情况:
- 大型语言模型 (LLM) 越来越多地被患者和临床医生使用.
- 虽然临床医生在管理患者信息等任务中使用LLM已有记录,但患者在理解临床笔记方面的利用仍未得到充分探索.
- 这项研究解决了了解患者如何利用LLM来解释健康信息的差距.
研究的目的:
- 评估商业上可用的LLM在响应患者生成的查询时的可靠性和准确性,基于开放的临床访问笔记.
- 为了比较不同LLM和提示策略在解释患者复杂的医疗信息的性能.
主要方法:
- 一项横截面的概念验证研究评估了三个LLM (ChatGPT 4o,Claude 3 Opus,Gemini 1.5).
- 四个提示系列 (标准,随机,人格,随机人格) 用患者设计的问题与神经瘤学进展说明相对应.
- 通过神经瘤学家和患者使用一个8个标准标题来评价答案,评估准确性,相关性,清晰性,可操作性,同理心,完整性,证据和一致性.
主要成果:
- 基于人格的提示,特别是ChatGPT 4o,在所有评估标准中获得了最高分.
- 标准和Persona提示系列的表现通常比Randomized或Randomized Persona系列的表现更好.
- 所有评估的LLM在提供证据以支持他们的反应方面表现不佳.
结论:
- LLM显示出有很大的潜力,可以帮助患者解释开放的临床笔记.
- 采用个性化的提示是优化患者驱动的健康信息查询中的LLM性能的一个关键策略.
- 进一步的开发和患者教育对于提高患者对健康数据的理解的LLM实用性至关重要.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


