Related Experiment Video
Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of Chunking and Embedding Strategies for Local Document Retrieval Using an Open-Source LLM in a Hospital
Jan Bossenz1, Carlo Günzl1, Fabian Berns2
1Junior Research Group (Bio-) Medical Data Science, Faculty of Medicine, Martin-Luther-University Halle-Wittenberg, Halle (Saale), Germany.
Introduction:
The integration of Retrieval-Augmented Generation (RAG) into domain-specific systems enables context-aware and traceable information retrieval. This study explores chunking and embedding strategies for a RAG-based question-answering system tailored to administrative documents at University Hospital Halle, focusing on model selection, parameter tuning, and retrieval performance. The insights gained from this study should serve as the foundation for the future development of a Retrieval-Augmented Generation (RAG) based chatbot system that aims to facilitate access to document pool contents for hospital staff.
Methods:
A corpus of 1,219 documents was preprocessed and chunked using varied parameters, including soft/hard character limits and overlaps. Eight embedding models were evaluated using Similarity Score and Maximum Marginal Relevance (MMR) retrievers. Top models Jinaai-v3 and Aari1995 were further analyzed across eight parameter configurations and ensemble retrievers using weight (w) and context (c) parameters.
Results:
Aari1995 reached the highest Top10 score (92.3%) with stable performance across chunk sizes and retriever configurations. Jinaai-v3 showed slightly stronger Top5 (84.6%) and Top3 (76.9%) scores but with greater sensitivity to parameter variations. Ensemble retrievers improved retrieval quality for both models, particularly when tuned via w-values. The c-parameter showed negligible influence. Runtime evaluation revealed that Jinaai-v3 generated vector stores more than four times faster than Aari1995. Overall, the similarity score retriever consistently outperformed MMR, both standalone and in ensemble configurations.
Conclusion:
Chunking and embedding choices significantly affect retrieval in domain-specific RAG systems. While both Jinaai-v3 and Aari1995 were effective, they differed in stability, accuracy, and efficiency. Findings support deploying a locally executable RAG system for administrative use, guiding future optimization of chunking and parameter robustness.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Chunking
The principle behind chunking...
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Methods of Documentation VII: EMR
Local Anesthetics: Clinical Application as Epidural Anesthesia
Since epidural anesthetics can be infused through an epidural catheter, all types of drugs, including short-acting ones, can be administered. Chloroprocaine and lidocaine are examples of short and long-duration anesthetics, respectively. Bupivacaine...
Local Anesthetics: Clinical Application as Spinal Anesthesia
Hospitals-II
Nurses that work in...