Related Experiment Video
Updated: Jun 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance and usability of retrieval-augmented large language models for stroke patient and caregiver support
Jinxia Rong1,2, Min Liang1, Zheyan Wang2
1School of Nursing, Fudan University, Shanghai, China.
Aim:
To develop a Retrieval-Augmented Generation (RAG) question-answering system for stroke patients and family caregivers, and evaluate its performance and usability.
Methods:
We constructed a localized knowledge base using clinical practice guidelines, consensus, expert opinion, textbooks, systematic review, evidence summary, and peer-reviewed literature. Three LLMs (GPT-4o, Claude3.7, and Qwen3) were first evaluated for accuracy using 184 exam questions under zero-shot and RAG configurations. Thirty open-ended stroke-related questions were assessed by three experienced clinicians across four dimensions. Usability testing was conducted with 20 stroke survivors and family caregivers using the best-performing model, measuring System Usability Scale (SUS) and Net Promoter Score (NPS).
Results:
RAG integration improved accuracy, relevance, and completeness across all three LLMs, with GPT-4o under RAG configuration achieving the highest overall mean score. However, the addition of RAG slightly reduced understandability for Claude3.7 and Qwen3. Usability testing yielded high acceptance.
Conclusions:
RAG can enhance the reliability of LLM-generated responses in stroke-related questions, offering trusted, guideline-based information. The high usability ratings suggest early feasibility for real-world deployment, while future research should assess its linguistic accessibility and long-term clinical benefits in real-world caregiving contexts.
