Related Experiment Video
Updated: Apr 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Real-world performance of open-source large language models in diabetes diagnosis
Shuting Yang1,2,3, Sujie Liu4,5, Yuxi Ma4,5
1National Clinical Research Center for Metabolic Diseases, Key Laboratory of Diabetes Immunology (Central South University), Ministry of Education, and Department of Metabolism and Endocrinology, The Second Xiangya Hospital of Central South University, Changsha, Hunan, China.
Background:
This study aimed to evaluate the performance of diverse open-source large language models (LLMs) in diagnosing diabetes subtypes and comorbidities from unstructured clinical text, assessing the impact of model characteristics, prompting, and language.
Methods:
We conducted a retrospective analysis of 11,329 adult diabetes patients from a large Chinese tertiary center (2010-2020). Various open-source LLMs were tested using four prompting strategies in English and Chinese. Primary outcomes were F1-scores for multi-class diabetes subtyping and binary classification of diabetic kidney disease (DKD) and metabolic syndrome (MetS).
Results:
LLMs demonstrated high performance in complex subtyping (peak F1 0.951) but showed limitations in rule-based DKD (F1 0.570) and MetS (F1 0.650) diagnosis. Chain-of-Thought prompting improved MetS classification but degraded DKD performance. Optimal model size was approximately 32B parameters. Notably, English prompts outperformed Chinese prompts on native Chinese text.
Conclusion:
Open-source LLMs exhibit strong holistic pattern recognition for complex classification but struggle with rule-based procedural reasoning. These models are promising as clinical co-pilots to augment expert decision-making rather than serving as autonomous diagnostic tools.
