Related Experiment Video
Updated: Sep 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Risks, Safety, and Reliability of Large Language Models in Military Medicine
Darshan Thota1, David Alt1, Sriram Venkatesan2
1Defense Health Agency, Falls Church, VA 22042, United States.
Objective:
To evaluate the performance of Large Language Models (LLMs) within the Military Health System (MHS), specifically assessing for bias, hallucinations, and safety risks.
Methods:
In late 2024, 47 clinicians across 13 military treatment facilities evaluated 8 blinded LLMs. Using standardized clinical vignettes, 835 conversations were reviewed for demographic bias, hallucinations, classification errors, and safety risks.
Results:
Demographic bias was present in 82 conversations (9.8%), hallucinations in 85 (10.2%), classification errors in 19 (2.3%), and safety risks in 78 (9.3%). Classification errors were linked to military-specific jargon. Safety risks were most prominent in overdose, critical care, and psychiatric scenarios.
Conclusion:
Failure to generate an accurate response can delay decision-making and increase cognitive burden and risk. Confidently incorrect responses are particularly dangerous. Although LLMs can generate useful clinical information, they remain vulnerable to biases and factual errors. Human oversight, training, and model tuning are necessary before utilizing military data.
