Related Experiment Video
Updated: Jan 28, 2026

Examining Bilingual Language Control Using the Stroop Task
Published on: February 26, 2020
Linguistic Disparities in Artificial Intelligence-Generated Patient Education for Total Hip Arthroplasty: A Pilot
Usman Ali1,2, Hafsa Khan Tareen3, Juan Antonio Pedroza4
1Department of Orthopedic Surgery, The Aga Khan University, Karachi, Pakistan.
Background:
Large Language Models (LLMs) are increasingly used for health information, but concerns exist regarding performance disparities for non-English speakers, potentially exacerbating health inequities. Appropriate information is critical for patients with limited English proficiency undergoing orthopedic procedures such as total hip arthroplasty (THA). This pilot study evaluated differences in the clinical reliability of English and Spanish responses to common THA questions generated by leading LLMs.
Methods:
Three widely accessible LLMs (ChatGPT-4o, Gemini 2.0 Flash, and Microsoft Copilot) were evaluated using 10 standardized frequently asked questions on THA, posed in English and Spanish. Responses were independently graded by language-fluent medical experts using a 4-point rubric (1 = Unsatisfactory to 4 = Excellent) assessing clinical reliability and appropriateness. Nonparametric statistics, including Wilcoxon signed-rank, Kruskal-Wallis, and effect sizes (Cliff's Delta, η2), were used for comparisons.
Results:
A statistically significant main effect of language was found (p = 0.014, η2 = 0.151), indicating significantly lower clinical reliability scores for Spanish responses in all LLMs. A nonsignificant within-model score decline was observed across all 3 LLMs.
Conclusion:
Leading LLMs exhibit significant difference in clinical reliability when providing THA information, performing less reliably in Spanish compared with English. This linguistic gap suggests a potential risk for difference in response interpretation and could potentially worsen health inequities for Spanish-speaking populations. Efforts are needed to improve multilingual capabilities and manage biases in medical artificial intelligence (AI). Clinicians and patients should exercise caution when using LLMs for health information in languages other than English until cross-lingual reliability is demonstrably improved.
Clinical Relevance:
This study highlights a significant linguistic disparity in AI-generated health information for THA. Improving LLMs' multilingual capabilities is essential to promote equitable access to reliable medical education and prevent the exacerbation of health inequities for non-English speaking patients.
Level Of Evidence:
Level IV. See Instructions for Authors for a complete description of levels of evidence.
1 To 2 Sentence Description:
This study evaluates LLMs in providing THA information in English and Spanish, revealing that Spanish responses are clinically less reliable. The findings highlight linguistic gap in AI healthcare tools, raising potential concerns for patient safety, and widening health inequities for non-English speakers.
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Cross-Sectional Research

