Related Experiment Videos
Multilingual Disparities in Large Language Model-Based Symptom Detection for Global Disease Surveillance: Evaluation
Sa'idah Zahrotul Jannah1,2, Tomohiro Nishiyama1, Shaowen Peng1
1Nara Institute of Science and Technology, 8916-5 Takayama-cho, Ikoma, Nara, 630-0192, Japan, 81 743725111.
Background:
Symptom detection is essential in global disease surveillance to detect potential outbreaks, as symptoms are the first observable signs of infection. To reflect real-time ground truth conditions during pandemics, social media has emerged as a valuable data source. Moreover, effective digital disease surveillance systems must operate across diverse linguistic settings, and large language models (LLMs) have been shown to perform inconsistently across languages, tending to have lower performance in low-resource languages. While multilingual approaches have been explored in various health-related natural language processing tasks, a critical gap remains in understanding whether LLM-based symptom detection can perform consistently across languages for global disease surveillance. Southeast Asia demonstrates this challenge, combining diverse languages and the potential for emerging infectious disease outbreaks, making it a case for evaluating how multilingual performance disparities manifest in symptom detection.
Objective:
This study aims to evaluate multilingual disparities in symptom detection using a LLM, as well as the associated error mechanisms across languages and symptom types, to better understand their implications for global disease surveillance.
Methods:
This study uses the MedWeb dataset, a multilingual pseudo-social media text dataset with multiple symptom labels. The data consist of 12 languages, covering diverse regions and language resource classifications. The symptoms included in this study are fever, headache, runny nose, cough, diarrhea, hay fever, influenza, and cold. We used GPT-5 as the symptom detection system, representing a strong model for health-related tasks. The results are evaluated using the macro precision, recall, and F1-score. An error analysis was conducted to identify the error mechanisms underlying incorrect symptom predictions across languages and to calculate the impact of each error category.
Results:
Our findings show that performance varied across languages, language resource groups, and symptom types. Southeast Asian (SEA) languages generally achieved lower scores than non-SEA languages, with Japanese obtaining the highest F1-score (0.812) and Lao the lowest (0.716). High-resource languages achieved the most consistent performance, while low-resource languages obtained the lowest overall scores. At the symptom level, diarrhea, headache, cough, and fever showed stable detection across languages, while hay fever, runny nose, influenza, and cold exhibited greater variability. Hay fever showed the widest variability, forming 2 distinct performance clusters aligned with language resource and region classifications. The error analysis revealed 4 misclassification patterns: explicitly mentioned symptoms, cross-lingual variation, symptom overgeneralization, and context misinterpretation. Cross-lingual variation was the most frequent error category, while errors related to explicitly mentioned symptoms showed the largest potential impact on model performance.
Conclusions:
Due to performance disparities, achieving more reliable and equitable LLM-based symptom detection from social media text for global disease surveillance would benefit from broader representation in training data for low-resource languages, improved cultural and linguistic sensitivity, and stronger contextual understanding of symptom-related expressions.