Related Experiment Video
Updated: Jun 13, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Language-Specific Differences in Large Language Model Diagnostic Reasoning: A Translation-Controlled Clinical
Jakub Magdziarz Ibrahim-El-Nur1, Wojciech Kaczmarek2, Weronika Winiarska1
1Department of Social Medicine and Public Health, Medical University of Warsaw, Pawińskiego 3a, 02-106 Warsaw, Poland.
Large language models (LLMs) show varied performance across languages in clinical diagnostics. Multilingual LLM evaluation is crucial before deploying them in non-English healthcare systems to ensure safety and efficacy.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Clinical Decision Support
Background:
- Large language models (LLMs) are increasingly assessed for clinical diagnostic tasks.
- LLM performance may differ significantly across languages, impacting their global applicability.
- Evaluating language influence is vital for safe deployment in diverse healthcare settings.
Purpose of the Study:
- To investigate if input language affects LLM diagnostic reasoning in clinical vignettes.
- To provide insights for multilingual predeployment evaluations of LLMs in non-English healthcare systems.
Main Methods:
- A translation-controlled in silico study using 30 clinical vignettes in English and Polish.
- Six LLMs were evaluated using a structured reflection framework with physician raters.
- Analysis of 720 rater-level evaluations and 360 model-language-vignette responses.
Main Results:
- Language significantly impacted LLM performance, varying by model; some models performed better in English.
- Differences were most pronounced in differential diagnosis quality and examination planning, not just final diagnosis.
- Language robustness was not consistent across evaluated LLMs, with performance gaps in safety-relevant reasoning domains.
Conclusions:
- Multilingual clinical performance of LLMs is highly model-dependent.
- Language-specific evaluations are essential before deploying LLMs in non-English healthcare systems.
- Performance disparities highlight the need for tailored validation to ensure clinical safety and effectiveness.
Related Concept Videos
Language and Cognition
Introduction to Language of Pathophysiology ll
Translation
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Translation
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Introduction to Language of Pathophysiology l
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...