Related Experiment Video
Updated: Feb 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Scaling Responsible Medical Text Retrieval Across Silos: Evaluating Small Language Models with Retrieval-Augmented
Rahul Shetty1, Karim Keshavjee1
1Institute of Health Policy, Management and Evaluation, Dalla Lana School of Public Health, University of Toronto, Toronto, ON, Canada.
None:
This study evaluates small language models with FAISS and Dynamic Top-k for diabetes-focused medical text retrieval, with all-mpnet-base-v2 performing best and FAISS (FAISS Facebook AI Similarity Search) maintaining sub-2 ms latency even at tenfold scale. Dynamic Top-k improved precision and nDCG (Normalized Discounted Cumulative Gain), showing that lightweight SLM-RAG (Small Language Model-Retrieval Augmented Generation) pipelines can approach larger-model accuracy while remaining highly scalable for cloud-based EMR (Electronic Medical Record) environments.
Related Concept Videos
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
pH Scale
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

