Related Experiment Video
Updated: Oct 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Semantic Clinical Knowledge Augmentation Improves Medical Question Answering Across Commercial and Open Large
Objective:
To evaluate whether semantic clinical knowledge augmentation through Semantic Clinical Artificial Intelligence (SCAI) improves medical question answering across commercial and open large language models, and to examine how a reasoning distilled model performs on the same benchmark.
Materials And Methods:
We evaluated Google Gemini, Microsoft Copilot, Meta Llama 3 70B, and DeepSeek-R1-Distill-Llama-70B on official text only United States Medical Licensing Exam (USMLE) sample questions from Step 1 (n = 87), Step 2 CK (n = 103), and Step 3 (n = 123), with and without SCAI. SCAI comprised HD-NLP parsing, graph and knowledge graph embeddings, a trained SCAI LLM semantic knowledge reasoner, and a semantic knowledge base; the current implementation contained 33,023,902 triples across 246 relation labels. Responses were scored against the answer key with secondary human verification. All models produced a response to every text only item, so incorrect responses were counted as confabulations.
Results:
Across 313 items, accuracy increased from 90.7% to 97.1% for Gemini, 82.4% to 92.0% for Llama 3 70B, and 55.3% to 81.8% for DeepSeek-R1-Distill; Copilot changed from 93.6% to 94.2%. Within model gains were statistically significant for Gemini, Llama 3 70B, and DeepSeek-R1-Distill on all three steps, but not for Copilot. The best single model result was Llama 3 70B plus SCAI on Step 3 (121/123, 98.4%). Gemini plus SCAI was the numerically strongest commercial system.
Conclusions:
SCAI was associated with materially better accuracy and lower confabulation (False Positive Results) across several heterogeneous LLMs, supporting a semantically grounded pipeline in which a dedicated SCAI LLM generates clinical context for a host LLM while suggesting that strong general reasoning performance does not automatically translate into healthcare readiness.
