Related Experiment Video
Updated: Oct 1, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Language-dependent variation in observed mechanistic performance of web-enabled large language models across
Jinxi Hu1, Tong Mei2, Lei Lin2
1Department of Orthopedics, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, China.
Objective:
To determine whether observed mechanistic performance of web-enabled large language model (LLM) responses differs across response languages under deployed web conditions and to distinguish such language-associated shifts in relative model performance from stable between-model differences.
Methods:
Twenty-four standardized questions spanning intestinal, airway, and urogenital mucosa and four levels of mechanistic complexity were posed to five web-enabled LLMs in English, Simplified Chinese, Traditional Chinese, and Japanese, yielding 480 responses. Twenty senior clinician-experts, assigned to language-specific panels and blinded to model identity, independently evaluated each response using a six-domain Mechanistic Fidelity Score (MFS; 0-24), with two ratings per response. Overall model effects were tested using block-adjusted regression with question-clustered inference. Because expert panels were nested within language, raw cross-language means were treated descriptively, whereas language-dependent relative performance was assessed after within-reviewer standardization. Each standardized value represents performance relative to an individual reviewer's own scoring distribution rather than an absolute cross-language MFS difference.
Results:
All 480 responses received two ratings (960 evaluations). Inter-rater agreement was moderate-to-good (Krippendorff α=0.709; Lin CCC = 0.709). Overall model performance differed substantially (F(4,23)=95.80, P<.001). Kimi K3 (mean MFS 20.65) and ChatGPT 5.6 Sol (20.43) outperformed DeepSeek V4 Flash 0731, ChatGPT 5.6 Luna, and Xiaomi V2.5 Pro Ultra after Holm correction (all P<.001), while differing nonsignificantly from each other (P = .182). Reviewer-standardized scores showed a strong model-by-language interaction in observed relative performance (F(12,191)=44.21, P<.001), whereas no statistically significant model-by-complexity interaction was detected in the present question set (P = .328).
Conclusion:
Mechanistic fidelity in the evaluated web-enabled systems was strongly model dependent, while observed relative performance varied across language panels. Because response language was not separable from panel composition and could also influence translation and web retrieval, these findings do not establish an intrinsic causal effect of language on underlying model reasoning. Multilingual biomedical evaluation should therefore test long-form mechanistic explanations directly in the language and deployment setting of intended use rather than infer equivalence from English performance alone.