Related Experiment Video
Updated: Oct 3, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Small language models in clinical medicine: a systematic review of performance, safety, and deployment feasibility
Alon Gorenshtein1,2,3, Mahmud Omar1, Yiftach Barash1,4
1BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA 02215, United States.
Objectives:
To review the clinical evidence for small language models (SLMs), 4 billion parameters or fewer, for performance, safety, and deployment feasibility.
Materials And Methods:
We searched 5 databases through February 10, 2026 for English-language reports of an SLM on a clinical task, following PRISMA 2020 under a PROSPERO-registered protocol. Two reviewers independently assessed eligibility (Cohen κ = 0.93) and rated methodological quality and transparency. We calculated a relative task score (RTS): the primary metric of each study's best-performing qualifying SLM, divided by a study-specific reference comparator selected by a uniform hierarchy, calculated by the review authors.
Results:
Eleven studies (7 peer-reviewed, 4 preprint or technical-report) were eligible. Across the 9 studies with a ratio-scale reference comparator, RTS ranged from 0.30 to 2.36; 1 was a raw difference on a non-ratio scale and 1 had no comparator. We did not pool estimates, given heterogeneity. Hallucination was evaluated in 6 of 11 studies, a scalar calibration-error metric or epistemic uncertainty in 0 of 11 (1 reported a calibration curve only), and on-hardware inference timing in 2 of 11. Our memory model estimated that a 4-billion-parameter model requires about 9.3 GB at a 2048-token context, indicating single-GPU memory feasibility rather than demonstrated deployment.
Discussion:
Domain-adapted SLMs are memory-feasible, but the evidence base is too small and heterogeneous for inferential comparison with larger models, and safety properties (including adversarial robustness) go unmeasured.
Conclusion:
Routine reporting of calibration and uncertainty and measurement of end-to-end clinical performance are prerequisites for responsible SLM deployment.
Prospero Registration:
CRD420261331444.