Related Experiment Video
Updated: Apr 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Real-Time Evaluation of a Large Language Model for Clinical Practice Guideline Development
1Department of Pharmacy Practice and Science, University of Arizona R. Ken Coit College of Pharmacy, Tucson, AZ.
Large language models (LLMs) show promise in framing healthcare questions for clinical practice guidelines but face limitations in later development stages. Further research is needed to optimize LLM capabilities in evidence synthesis and decision-making frameworks.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Practice Guideline Development
- Health Informatics
Background:
- Clinical practice guideline development is a complex, multi-step process.
- Large language models (LLMs) are emerging tools with potential applications in healthcare.
- Evaluating LLM capabilities in guideline development is crucial for understanding their utility.
Purpose of the Study:
- To assess the effectiveness of a large language model (LLM) in executing all phases of clinical practice guideline development.
- To evaluate the performance of OpenAI's Generative Pretrained Transformer (GPT)-4o in guideline creation.
Main Methods:
- The study utilized GPT-4o to perform tasks aligned with the Grading of Recommendations, Assessment, Development, and Evaluation (GRADE) framework.
- LLM prompts were executed concurrently with the guideline panel's progress on a guideline for neuromuscular blockade in acute respiratory distress syndrome.
- The evaluation covered question framing, outcome selection/rating, evidence summarization, quality assessment, and evidence-to-decision frameworks.
Main Results:
- The LLM demonstrated significant utility in the initial stages, particularly in framing healthcare questions and selecting/rating outcomes.
- The LLM's limitations became more evident in subsequent steps, including evidence summarization and quality assessment.
- The evidence-to-decision framework creation highlighted the LLM's challenges in synthesizing complex evidence.
Conclusions:
- LLMs show the most promise in the early, foundational steps of clinical practice guideline development.
- Current LLM capabilities are limited in the more intricate phases of evidence synthesis and recommendation formulation.
- Further development and validation are required to enhance LLM performance across the entire guideline development lifecycle.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Introduction to Language of Pathophysiology ll
Improving Translational Accuracy
Improving Translational Accuracy
Introduction to Language of Pathophysiology l
Nursing Evaluation