Related Experiment Video
Updated: Jul 2, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs
Li Wang1,2, Xi Chen1,2, XiangWen Deng3
1Sports Medicine Center, West China Hospital, Sichuan University, Chengdu, China.
Prompt engineering enhances large language model (LLM) accuracy in clinical medicine. Specific prompts, like ROT for gpt-4-Web, improve consistency with orthopedic guidelines.
Area of Science:
- Clinical Medicine
- Artificial Intelligence
- Medical Informatics
Background:
- Large language models (LLMs) show promise in clinical medicine.
- Effective knowledge transfer from computer science to clinical applications is essential.
- Prompt engineering is a key method for optimizing LLM performance.
Purpose of the Study:
- To evaluate the impact of prompt engineering on LLM reliability and accuracy in clinical medicine.
- To assess LLM agreement with the American Academy of Orthopedic Surgeons (AAOS) osteoarthritis (OA) guidelines.
- To compare the consistency of different prompts and LLMs.
Main Methods:
- Designed and applied various prompt styles to different LLMs.
- Queried LLMs regarding adherence to AAOS OA evidence-based guidelines.
- Repeated each query five times to assess reliability.
- Analyzed consistency across different evidence levels and prompt types.
Main Results:
- GPT-4-Web with ROT prompting achieved the highest overall consistency (62.9%).
- This prompt style demonstrated strong performance for significant recommendations (77.5% consistency).
- LLM reliability varied significantly across prompts and models (Fleiss kappa: -0.002 to 0.984).
Conclusions:
- Prompt engineering significantly influences LLM performance in clinical contexts.
- The ROT prompt for GPT-4-Web emerged as the most consistent method.
- Careful prompt selection can enhance the accuracy of LLM responses to medical queries.
More Related Videos
12:55Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
Published on: September 27, 2020
13:05Reliable Mechanochemistry: Protocols for Reproducible Outcomes of Neat and Liquid Assisted Ball-mill Grinding Experiments
Published on: January 23, 2018
Related Concept Videos
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Design Example: Managing Concrete Workability
Legal Guidelines for Documentation
Data Validation
Key parameters for method validation include:
Design Example: Maintaining Level of an Embankment
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...