快速工程在一致性和可靠性与LLMs的基于证据的指南
Li Wang1,2, Xi Chen1,2, XiangWen Deng3
1Sports Medicine Center, West China Hospital, Sichuan University, Chengdu, China.
快速工程提高了临床医学中大语言模型 (LLM) 的准确性. 特定提示,如gpt-4-Web的ROT,提高了与骨科指南的一致性.
科学领域:
- 临床医学 临床医学
- 人工智能的人工智能
- 医疗信息学 医疗信息学
背景情况:
- 大型语言模型 (LLM) 在临床医学中表现有前途.
- 从计算机科学到临床应用,有效的知识转移至关重要.
- 提示工程是优化LLM性能的一个关键方法.
研究的目的:
- 评估快速工程对临床医学LLM可靠性和准确性的影响.
- 评估与美国骨科外科医生学会 (AAOS) 骨关节炎 (OA) 准则的LLM协议.
- 为了比较不同提示和LLMs的一致性.
主要方法:
- 设计并应用各种提示风格到不同的LLMs.
- 查询了关于遵守AAOS OA基于证据的指导方针的LLMs.
- 重复每个查询五次以评估可靠性.
- 分析了不同证据级别和提示类型的一致性.
主要成果:
- 带有ROT提示的GPT-4-Web实现了最高的整体一致性 (62.9%).
- 这种快速的风格显示了重要的建议 (77.5%的一致性) 的强表现.
- 在各种提示和模型中,LLM的可靠性差异很大 (Fleiss kappa: -0.002到0.984).
结论:
- 快速工程在临床环境中显著影响LLM业绩.
- 对GPT-4-Web的ROT提示符成为最一致的方法.
- 仔细的快速选择可以提高LLM对医疗查询的答案的准确性.
更多相关视频
12:55Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
Published on: September 27, 2020
13:05Reliable Mechanochemistry: Protocols for Reproducible Outcomes of Neat and Liquid Assisted Ball-mill Grinding Experiments
Published on: January 23, 2018
相关概念视频
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Design Example: Managing Concrete Workability
Legal Guidelines for Documentation
Data Validation
Key parameters for method validation include:
Design Example: Maintaining Level of an Embankment
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
