Related Experiment Video
Updated: Jun 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models on multilingual vaccine knowledge: a benchmark study
Siyuan Chen1,2, Lily Wass1,2, Zhengdong Wu1,2,3
1Laboratory of Data Discovery for Health Limited (D24H), Hong Kong Science Park, Hong Kong SAR, China.
None:
Large language models (LLMs) are increasingly used by clinicians and the public for vaccine information, yet their factual accuracy across languages and vaccine domains remains insufficiently characterized. We evaluated 13 LLMs using VaxEval, a multilingual vaccine-knowledge benchmark of 1886 vaccine-related multiple-choice questions spanning 14 vaccines in English (71%), Spanish (13%), and Chinese (16%). All items underwent quality control, with reference answers verified against authoritative guidance and peer-reviewed sources. Model performance was evaluated under zero-shot, few-shot, and chain-of-thought (CoT) prompting, with exact-match accuracy defined as selecting the pre-specified reference option. We used mixed-effect logistic regression to estimate associations between model group (newer flagship models vs earlier models), prompting strategy, language, and vaccine type, and answer correctness. Mean accuracy across models was 86.0% in English, 83.7% in Spanish, and 80.0% in Chinese. Flagship models had higher odds of correctness than earlier versions (OR 1.57; 95% CI 1.50-1.65; P < .001). Few-shot prompting was associated with higher correctness (OR 1.17; P < .001), whereas CoT prompting was associated with lower correctness (OR 0.79; P < .001). Performance varied by vaccine type and question category, underscoring the need for rigorous evaluation, structured guardrails, and targeted refinement before using LLMs for vaccine communication.
Related Concept Videos
Vaccines
Vaccinations
Improving Translational Accuracy
Improving Translational Accuracy
