Related Experiment Video
Updated: Jan 8, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Artificial intelligence-driven clinical guideline recommendations in maternal care: How trustworthy are they?
Jairo J Pérez1, Andrés F Giraldo-Forero2, Santiago Rúa3
1Departamento de Ciencias Aplicadas, Instituto Tecnológico Metropolitano, Medellín, Colombia.
Introduction:
Medical staff often face difficulties in consulting and applying clinical guidelines in practice. Large language models, especially when combined with retrieval-augmented generation, may help overcome these challenges by producing context-specific outputs with improved adherence to medical guidelines.
Objectives:
To assess the performance of commercial large language models in answering maternal health questions within retrieval-augmented generation systems, using both human and automated evaluation metrics.
Material And Methods:
A controlled experiment was designed to obtain accurate, consistent answers from a retrieval-augmented generation system based on Colombian maternal care guidelines. A physician formulated ten questions and defined the groundtruth answers. Various large language models were tested with a standardized prompt and evaluated through binary answer-concept ranking and retrieval-augmented generation assessment, metrics, judged by two independent large language models.
Results:
Generative pre-trained transformer 3.5 (GPT-3.5) achieved the highest physicianassessed accuracy (0.90). Claude 3.5 obtained the top faithfulness score (0.78) under GPT-4.o evaluation, while Mistral ranked highest (0.84) under Claude 3.5 evaluation. Regarding answer relevance, GPT-3.5 scored highest across both judges (0.94 and 0.86).
Conclusions:
Integrating retrieval-augmented generation into obstetric care has the potential to enhance evidence-based practices and improve patient outcomes. However, rigorous validation of accuracy and context-specific reliability is essential before clinical deployment. The findings of this study indicate that large-scale models (e.g., GPT-3.5, Claude, Llama 70B) consistently outperform lighter models such as Llama 8B.
Related Concept Videos
Current Trends in Nursing II
Ethical Issues
Ethical Concerns in Healthcare:
Ethical Dilemmas I
Let us explore some examples to understand the potentially complex moral decisions nurses face.
Take the case of caring for minors, particularly in areas related to reproductive...
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Standards of Care II
Standards of Care I