Related Experiment Video
Updated: Jul 4, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploratory Evaluation of Large Language Models for Reducing Language Bias in Systematic Review Screening
Junki Ikeguchi1, Hiroaki Ueshima2,3, Hiroshi Tamura1,2
1Graduate School of Informatics, Kyoto University, Kyoto, Japan.
Abstract:
Language bias arises in systematic reviews when non-English studies are excluded owing to resource constraints. Large language models (LLMs) can mitigate this problem through multilingual processing. To assess whether direct multilingual LLM processing reduces language-based disparities in systematic review screening performance compared to translation-mediated approaches. Six state-of-the-art LLMs were evaluated under three conditions: (1) an English benchmark dataset (n = 2,911), (2) direct screening of non-English abstracts (n = 483), and (3) screening of machine-translated non-English abstracts. Performance was measured using sensitivity, specificity, F1 score, balanced accuracy, and workload reduction. All models achieved high sensitivity on English data (≥0.938). Translation-mediated screening substantially reduced sensitivity in some models (range: 0.47-0.54), whereas direct multilingual processing maintained high sensitivity (range: 0.71-1.00). Considerable differences were observed among models. Direct multilingual LLM screening may reduce language-related sensitivity disparities; however, the effects on downstream meta-analytic bias require further investigation.
