Related Experiment Video
Updated: Sep 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective
Haiping Zhang1,2, Xinming Li3,4, Kechi Fang3,4
1Clinical Laboratory Center, Beijing Youan Hospital, Capital Medical University, Beijing, P R China.
Background:
Differentiating among liver disease entities such as autoimmune liver disease (AILD), drug-induced liver injury (DILI), and chronic hepatitis B (CHB) remains clinically challenging due to overlapping clinical manifestations and nonspecific laboratory findings. Conventional machine learning (ML) approaches rely mainly on structured laboratory data, whereas free-text clinical reports and other heterogeneous electronic medical record data are often underused. Large language models (LLMs) may provide a strategy for encoding heterogeneous clinical information, yet their usefulness for liver disease classification remains insufficiently evaluated.
Objective:
This study aimed to evaluate the usefulness of LLM-derived embeddings for clinical data mining in liver disease and to determine whether integrating these embeddings with laboratory variables improves classification across broad disease categories and closely related subtypes.
Methods:
We retrospectively analyzed electronic medical record data from 7543 patients with nonoverlapping liver disease etiologies treated at Beijing Youan Hospital, Capital Medical University, between 2010 and 2025. Three LLMs (Qwen3, Huatuo-o1, and II-Medical) generated semantic embeddings from standardized clinical text, combining free-text examination reports, and structured clinical observations. Performance was assessed in a 3-class etiological task (AILD, DILI, and CHB) and a 4-class task further subclassifying AILD into autoimmune hepatitis and primary biliary cholangitis. We compared embedding-only models, LLM-integrated ML models, and an ML-only baseline using the same structured variable set and preprocessing pipeline, with lightweight natural language processing encoders and zero-shot LLM reasoning as additional comparators. Models were developed using 5-fold cross-validation and evaluated on an internal holdout set using accuracy, macroaveraged precision, recall, and F1-score.
Results:
In the 3-class task, the LLM-integrated ML models achieved macro F1-scores of 0.835-0.837, compared with 0.791 for the ML-only baseline, with corresponding accuracies of 0.925-0.929 versus 0.893. In the 4-class task, the LLM-integrated ML models achieved macro F1-scores of 0.717-0.734, compared with 0.665 for the ML-only baseline, with corresponding accuracies of 0.920-0.922 versus 0.874. A temporal split sensitivity analysis using cases from 2010 to 2019 for training and cases from 2020 to 2025 for testing showed that the relative advantage of LLM-integrated ML models over the ML-only baseline was preserved. Direct zero-shot LLM reasoning and lightweight natural language processing encoders performed below the embedding-based integrated models.
Conclusions:
In this single-center retrospective cohort of patients with clear-cut, nonoverlapping liver disease etiologies, LLM-derived embeddings provided complementary information to structured laboratory variables for multiclass liver disease classification. The integrated framework showed improved internal validation performance compared with the ML-only model, particularly for non-CHB categories and fine-grained subtype discrimination. Because patients with overlapping liver disease etiologies were excluded, the reported performance may overestimate diagnostic accuracy in broader real-world clinical settings where overlapping syndromes are common. Multicenter external validation and prospective evaluation in more heterogeneous patient populations are needed before clinical implementation.