Related Experiment Video
Updated: Apr 10, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
ClinicRealm: Re-evaluating large language models with conventional machine learning for non-generative clinical
Yinghao Zhu1,2,3, Junyi Gao4,5, Zixiang Wang2
1School of Artificial Intelligence, Beihang University, Beijing, China.
NPJ Digital Medicine
|April 8, 2026
Summary
Modern Large Language Models (LLMs) now outperform specialized models for clinical prediction using unstructured notes. Advanced LLMs also show strong performance on structured Electronic Health Records (EHR), even with limited data.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Machine Learning for Healthcare
Background:
- Large Language Models (LLMs) are increasingly used in healthcare, but their effectiveness for non-generative clinical prediction tasks remains under-evaluated.
- There is a common assumption that specialized models are superior to LLMs for clinical prediction, potentially leading to misuse and misunderstanding.
- The ClinicRealm benchmark was developed to systematically assess LLM performance in clinical settings.
Purpose of the Study:
- To systematically evaluate the performance of various Large Language Models (LLMs) against traditional methods for clinical prediction tasks.
- To compare LLM capabilities on both unstructured clinical notes and structured Electronic Health Records (EHR).
- To assess predictive performance, reasoning abilities, and fairness across different model types.
Main Methods:
- Evaluated 15 GPT-style LLMs, 5 BERT-style models, and 11 traditional methods using the ClinicRealm benchmark.
- Utilized unstructured clinical notes and structured Electronic Health Records (EHR) data for evaluation.
- Assessed models on predictive performance, reasoning, and fairness metrics.
Main Results:
- Leading zero-shot LLMs (e.g., DeepSeek-V3.1-Think, GPT-5) significantly outperformed fine-tuned BERT models on clinical notes.
- On structured EHR data, advanced LLMs demonstrated strong zero-shot capabilities, often surpassing conventional models in data-scarce scenarios.
- Leading open-source LLMs achieved performance comparable to or exceeding proprietary models.
Conclusions:
- Modern LLMs are highly competitive for clinical prediction tasks, challenging the assumption of their inferiority to specialized models.
- LLMs show significant promise for both unstructured and structured clinical data, particularly in data-limited situations.
- Health data scientists and developers should re-evaluate model selection strategies considering the advanced capabilities of contemporary LLMs.
Related Concept Videos
Improving Translational Accuracy
15.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.6K
Improving Translational Accuracy
3.8K
3.8K
