Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in
Yuhan Zhao1, Yiqun Miao1, Rongrong Guo1
1School of Nursing, Capital Medical University, No. 10 Xitoutiao, Youanmenwai, Fengtai District, Beijing, 100069, China, 86 13910789837.
Journal of Medical Internet Research
|August 3, 2026
Summary
Evaluation of retrieval-augmented generation (RAG) and graph-structured RAG (GraphRAG) in healthcare LLMs is inconsistent. Gaps exist in real-world testing, safety, and detailed verification, hindering clinical readiness.
Area of Science:
- * Artificial Intelligence in Healthcare
- * Natural Language Processing
- * Clinical Informatics
Background:
- * Retrieval augmentation is crucial for traceable and verifiable large language model (LLM) applications in healthcare.
- * Current evaluation practices for text-based retrieval-augmented generation (RAG) and graph-structured RAG (GraphRAG) are heterogeneous, impeding cross-study comparisons and clinical readiness assessments.
Purpose of the Study:
- * To map and characterize evaluation methods for inference-time RAG and GraphRAG systems in healthcare.
- * To analyze how evaluation constructs are defined, operationalized, and reported across different system layers and settings.
Main Methods:
- * Conducted a scoping review following PRISMA-ScR guidelines, searching multiple databases up to May 2026.
- * Included studies on healthcare-relevant LLM systems using inference-time RAG with at least one evaluation component.
- * Extracted data on study characteristics, system design, retrieval, evidence linkage, safety, GraphRAG specifics, reporting, and governance.
Main Results:
- * 157 studies were included, primarily focusing on clinical question answering and decision support.
- * Most evaluations were offline (89.2%), with limited workflow-facing or prospective assessments.
- * Significant gaps were identified in evaluating retrieval quality, fine-grained evidence verification, safety, LLM-as-judge bias, and GraphRAG components in real-world settings.
Conclusions:
- * Healthcare RAG and GraphRAG system evaluation is rapidly growing but lacks consistent reporting and definitions.
- * A shift towards more transparent, layer-specific, safety-oriented, and clinically contextualized evaluations is needed.
- * Addressing identified gaps is critical for advancing the clinical implementation of these AI systems.
