Related Experiment Video
Updated: Jun 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
How Reliably Do Large Language Models Reproduce Vital Pulp Therapy Guidelines? A Mixed-Effects Evaluation of
Sine Güngör Us1, Arzu Şahin Mantı2, Arzu Kaya Mumcu3
1Department of Endodontics, Faculty of Dentistry, Gazi University, 06490 Ankara, Turkey.
Abstract:
Background: Large language models (LLMs) are increasingly consulted for clinical guidance, yet their reliability in protocol-sensitive domains remains insufficiently characterized. This study evaluated the ability of widely accessible LLMs to reproduce guideline-defined decision thresholds in vital pulp therapy (VPT), with emphasis on guideline-concordance accuracy, professional-role prompting, short-term response stability, and decision-level error directionality. Methods: Twenty-six binary yes/no questions were derived from an internationally recognized evidence-based guideline for VPT. Four LLMs-GPT-5, GPT-4o, DeepSeek-V3, and Gemini 2.5 Flash-were queried under non-prompted and professional-role-prompted conditions by two independent operators across three daily sessions over three consecutive days. Descriptive analyses were complemented by mixed-effects logistic regression in R to account for repeated responses clustered within guideline-derived questions. Results: Overall guideline-concordance accuracy was high across models. Gemini showed the highest observed accuracy under non-prompted conditions; DeepSeek showed the highest under prompted conditions. In the mixed-effects model, Gemini demonstrated significantly higher odds of guideline-concordant responses than GPT-5 under non-prompted conditions, whereas DeepSeek outperformed GPT-5 and GPT-4o under prompted conditions. The model × prompt interaction showed a trend toward significance but did not reach the conventional threshold. Day and within-day time point were not significantly associated with accuracy, supporting short-term response stability. Error-direction analysis revealed model-specific patterns: Gemini showed consistently low false-positive rates but increased false-negative responses under prompted conditions; DeepSeek showed reduced false-positive and no false-negative responses under prompted conditions. Conclusions: Average accuracy alone is insufficient to characterize the reliability of LLM-generated clinical guidance. Evaluation in protocol-sensitive domains should incorporate guideline-concordance, prompt responsiveness, short-term stability, and decision-level error directionality.
Related Concept Videos
Guidelines For Measuring Vital Signs
Before taking a patient's vital signs, a nurse would consider and assess the patient's comfort level and ensure appropriate equipment is available.
Pre-Procedural Guidelines for Assessing Blood Pressure
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care evaluation by...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Guidelines for Nursing Documentation II
Timely documentation is crucial to ensure continuity of care for patients. Any delays in recording or reporting medical information can result in medical errors and even adverse patient outcomes. From medication administration to diagnostic test results, every detail must be accurately and promptly documented to provide the best possible care for patients.