Related Experiment Video
Updated: Apr 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Real-world performance of open-source large language models in diabetes diagnosis
Shuting Yang1,2,3, Sujie Liu4,5, Yuxi Ma4,5
1National Clinical Research Center for Metabolic Diseases, Key Laboratory of Diabetes Immunology (Central South University), Ministry of Education, and Department of Metabolism and Endocrinology, The Second Xiangya Hospital of Central South University, Changsha, Hunan, China.
Open-source large language models (LLMs) excel at complex diabetes subtyping but struggle with rule-based diagnoses like diabetic kidney disease. English prompts performed better on Chinese text, suggesting LLMs can aid clinicians but not replace them.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing in Healthcare
- Clinical Decision Support Systems
Background:
- Evaluating open-source large language models (LLMs) for diagnosing diabetes subtypes and comorbidities from clinical text.
- Assessing the influence of LLM characteristics, prompting techniques, and language on diagnostic performance.
- Focus on unstructured clinical data from a large patient cohort.
Purpose of the Study:
- To assess the diagnostic capabilities of various open-source LLMs for diabetes-related conditions.
- To investigate the impact of different prompting strategies and model parameters on performance.
- To compare the effectiveness of English versus Chinese prompts for analyzing Chinese clinical text.
Main Methods:
- Retrospective analysis of 11,329 adult diabetes patients (2010-2020).
- Testing diverse open-source LLMs with four prompting strategies in English and Chinese.
- Primary outcome metrics: F1-scores for diabetes subtyping, diabetic kidney disease (DKD), and metabolic syndrome (MetS) classification.
Main Results:
- LLMs achieved high performance in complex diabetes subtyping (peak F1 0.951).
- Performance was limited in rule-based diagnoses: DKD (F1 0.570) and MetS (F1 0.650).
- English prompts outperformed Chinese prompts on native Chinese text; Chain-of-Thought prompting showed mixed results.
Conclusions:
- Open-source LLMs demonstrate strong pattern recognition for complex classification tasks.
- LLMs exhibit limitations in rule-based procedural reasoning, impacting specific diagnoses.
- These models show potential as clinical co-pilots to support, not replace, expert medical decision-making.
