Related Experiment Video
Updated: Jun 21, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Large Language Models for Postoperative Decision Support: A
Srinivasagam Prabha1, Bernardo Gabriele Collaco1, Cesar Abraham Gomez-Cabello1
1Division of Plastic Surgery, Mayo Clinic in Florida, 4500 San Pablo Rd S, Jacksonville, US.
Journal of Medical Internet Research
|June 19, 2026
Summary
Knowledge-enhanced large language models (LLMs) significantly improve postoperative decision support accuracy. The hybrid fine-tuning (FT) plus retrieval-augmented generation (RAG) approach demonstrated the highest performance, showing promise for patient education.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Natural Language Processing
Background:
- Large language models (LLMs) offer potential for clinical decision support but struggle with integrating domain-specific medical knowledge for tasks like postoperative patient education.
- Challenges include maintaining accuracy, safety, and interpretability when adapting LLMs for healthcare applications.
- Fine-tuning (FT), retrieval-augmented generation (RAG), and hybrid FT+RAG are key strategies for knowledge integration, yet their comparative efficacy in postoperative care is unevaluated.
Purpose of the Study:
- To compare the performance, reliability, and safety of baseline, FT, RAG, and hybrid FT+RAG LLM configurations for postoperative decision support.
- To evaluate LLM accuracy in routine and emergency postoperative scenarios.
- To assess the impact of knowledge integration strategies on LLM performance in a clinical context.
Main Methods:
- A comparative evaluation of four LLM configurations (baseline, FT, RAG, FT+RAG) using Google Gemini 2.5 Flash.
- Model adaptation and validation using 600 postoperative question-answer pairs, with final evaluation on 150 queries including routine, emergency, and out-of-scope prompts.
- Independent assessment by 3 blinded clinical experts for accuracy, safety, completeness, and relevance, supplemented by automated metrics for readability and hallucination.
Main Results:
- All knowledge-enhanced LLMs significantly outperformed the baseline model in overall accuracy (FT: 92.7%, RAG: 91.3%, FT+RAG: 97.3% vs. baseline: 68.0%).
- The FT+RAG configuration achieved the highest clinical medical accuracy (96.7%) for in-scope queries and demonstrated superior composite classification performance (100% precision, 96.7% recall, 98.3% F1 score).
- Knowledge-enhanced models showed improved safety/refusal accuracy, though baseline comparisons were influenced by differing safety instructions; readability was generally lower due to safety boilerplate.
Conclusions:
- Incorporating domain-specific knowledge via FT, RAG, or FT+RAG substantially improves LLM performance for postoperative decision support compared to baseline models.
- The hybrid FT+RAG approach yielded the most favorable outcomes, indicating its potential for enhancing postoperative patient education and decision-making.
- Further validation, readability optimization, and robust governance are essential before widespread patient-facing deployment of knowledge-enhanced LLMs in healthcare.