Related Experiment Video
Updated: Mar 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance Evaluation of Large Language Models in Multilingual Medical Multiple-Choice Questions: Mixed Methods
Livia Maria Strasser1, Wilma Anschuetz2, Fabio Dennstädt1,3
1Medical Knowledge and Decision Support, School of Medicine, University of St.Gallen, St.Jakobstrasse 21, St.Gallen, 9000, Switzerland, +41 71 224 32 00.
Large language models (LLMs) show varied accuracy on medical questions across languages, with German performing best. Prompting in English generally improved results, but human oversight is crucial for reliable integration into medical education.
Area of Science:
- Medical Education
- Artificial Intelligence
- Natural Language Processing
Background:
- Artificial intelligence (AI) is transforming healthcare and medical education.
- Large language models (LLMs) demonstrate potential in medical licensing exams.
- LLM performance varies by language, necessitating cross-language comparisons.
Purpose of the Study:
- Evaluate LLM performance on medical multiple-choice questions across German, French, and Italian.
- Quantitatively and qualitatively assess LLM capabilities in multilingual medical education.
- Identify factors influencing LLM accuracy in diverse linguistic contexts.
Main Methods:
- Mixed methods study analyzing 114 multiple-choice questions in German, French, and Italian.
- Quantitative performance analysis of multiple LLMs (OpenAI, Meta AI, Anthropic, DeepSeek).
- Qualitative analysis of answer explanations from top-performing LLMs (GPT4o, Claude-Sonnet-3.7) for incorrect answers.
Main Results:
- LLM accuracy varied significantly by model and language (64%-87%), with German questions yielding the best performance.
- English prompts generally outperformed language-matched prompts, though top models showed comparable results.
- Qualitative analysis revealed reasoning errors in LLM explanations and identified 3 imprecise questions.
Conclusions:
- LLM performance in medical exams is influenced by model, prompt, and input language, requiring careful selection.
- LLM-generated explanations can enhance medical question quality, contingent on data security.
- Human oversight is essential for nuanced medical content, and ongoing evaluation is needed for reliable LLM integration.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Related Concept Videos
Language and Cognition
Improving Translational Accuracy
Improving Translational Accuracy