快速工程和后续询问提高了大型语言模型中脊柱手术问题的可读性
Sohail Daulat1, Nikhil Dholaria2, Gregory Burnet1
1Department of Neurological Surgery, University of Pittsburgh School of Medicine, Pittsburgh, PA.
World neurosurgery
|September 1, 2025
概括
快速的工程和后续问题显著提高了脊椎手术患者教育的大型语言模型 (LLM) 的可读性. 聊天GPT-4o通常比聊天GPT-5更容易阅读的内容.
科学领域:
- 医学教育
- 医疗保健中的人工智能
- 脊椎手术
背景情况:
- 在脊椎外科的患者教育材料往往超过建议的阅读水平.
- 大型语言模型 (LLM) 对产生教育内容有希望,但需要进一步研究可读性和适应性.
- 评估LLM在创建可访问的患者信息方面的表现至关重要.
研究的目的:
- 评估哪个LLM模型和提示策略最能提高脊椎手术患者教育的可读性.
- 为了比较新的LLM模型在生成可理解的内容方面的表现.
- 确定最优的方法来改进LLM产生的教育材料.
主要方法:
- 查特GPT-4o和查特GPT-5得到了五种常见手术中的45个脊椎手术问题.
- 通过五个提示阶段生成响应,包括基线,后续澄清,六年级级要求,基于规则的提示和直接可读性定位.
- 使用多种评分系统 (SMOG,FRE,FKGL,GFI,CLI) 评估可读性,并进行统计分析.
主要成果:
- 在大多数阶段,ChatGPT-4o产生的反应明显比ChatGPT-5更容易读取 (p<0. 001).
- 六年级级的要求阶段 (第二阶段) 得到了最易读的答案,超过51%的答案达到目标.
- 后续澄清和简化提示比基于规则的复杂策略更有效地提高可读性.
结论:
- 快速的工程和后续询问大大提高了患者对LLM产生的脊椎手术内容的可读性.
- 虽然ChatGPT-4o显示出更好的可读性,但ChatGPT-5提供了更可靠的引用.
- 未来的研究应该在客观评分之外的现实临床环境中验证这些发现.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


