Related Experiment Video
Updated: Sep 13, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Development and evaluation of an agentic LLM based RAG framework for evidence-based patient education
AlHasan AlSammarraie1, Ali Al-Saifi2, Hassan Kamhia3
1Hamad Bin Khalifa University College of Science and Engineering, Doha, Qatar aalsammarraie@hbku.edu.qa.
Agentic retrieval augmented generation (ARAG) improves Arabic patient education material (PEM) generation using large language models (LLMs). Larger models are crucial for validating and blocking harmful content effectively.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Health Informatics
Background:
- Developing high-quality, evidence-based patient education materials (PEMs) is crucial for effective healthcare communication.
- Large language models (LLMs) offer potential for automating content generation, but require careful evaluation for accuracy and safety, especially in specialized languages like Arabic.
Purpose of the Study:
- To develop and evaluate an agentic retrieval augmented generation (ARAG) framework for generating Arabic PEMs using open-source LLMs.
- To assess LLMs' capability to act as validation agents for blocking harmful content within generated PEMs.
Main Methods:
- Twelve LLMs were tested across four setups: base, base with prompt engineering, ARAG, and ARAG with prompt engineering.
- PEM quality was evaluated using automated LLM assessment and expert review across five metrics: accuracy, readability, comprehensiveness, appropriateness, and safety.
- Validation agent performance was assessed using a dataset of harmful/safe PEMs to measure blocking accuracy.
Main Results:
- ARAG setups enhanced generation performance for 10 out of 12 LLMs, with Arabic-focused models dominating the top ranks.
- AceGPT-v2-32B with ARAG and prompt engineering demonstrated the highest performance.
- Validation agent accuracy strongly correlated with model size (≥27B parameters needed for >0.80 accuracy), though Fanar-7B excelled in generation but not validation.
Conclusions:
- Arabic-centric LLMs show advantages for Arabic PEM generation.
- The ARAG framework significantly improves PEM generation quality, though context limitations affect large-context models.
- Model size is a critical factor for reliable harmful content validation by LLMs.
More Related Videos
Related Concept Videos
Nursing Process for Patient and Caregiver Teaching III: Evaluation and Documentation
Nurses can use several methods to evaluate patient outcomes. For example, oral questions can assess cognitive learning,...
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Health Literacy
Role of Communication in the Nursing Process III: Evaluation and Documentation
Nursing Process for Patient and Caregiver Teaching I: Assessment and Diagnosis
It is critical to determine the patient's learning needs during the assessment. Determination of learning needs compounds data...
Methods of Documentation VII: EMR

