Related Experiment Video
Updated: Apr 4, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Diagnostic Accuracy of Large Language Models for Rare Diseases: A Systematic Review and Meta-Analysis
Minh-Ha Nguyen1, Chih-Ting Yang2, Thomas A Cassini3
1Department of Epidemiology, Vanderbilt University, Nashville, TN, USA.
Medrxiv : the Preprint Server for Health Sciences
|April 3, 2026
Summary
Large language models (LLMs) show promise for rare disease diagnosis, but performance varies. Augmented LLMs and benchmarks with fewer ultra-rare diseases performed better, though bias remains a concern.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Rare Disease Diagnostics
Background:
- Evidence on large language models (LLMs) for rare disease diagnosis is fragmented.
- LLMs show potential but require rigorous evaluation for clinical use.
- Systematic review and meta-analysis to assess LLM diagnostic accuracy for rare diseases.
Purpose of the Study:
- Synthesize evidence on LLM diagnostic performance in rare diseases.
- Identify factors influencing LLM accuracy (e.g., knowledge augmentation, input modality).
- Evaluate the readiness of LLM technology for clinical translation.
Main Methods:
- Systematic search of major databases (PubMed, Embase, etc.) and preprint servers.
- Inclusion of studies evaluating LLM-based rare disease differential diagnosis.
- Meta-analysis of Recall@1 (R@1) using random-effects models; subgroup and post-hoc analyses.
Main Results:
- Pooled R@1 was 43.3% across 19 system-dataset entries (N=39,529 cases).
- Augmented LLMs (52.5% R@1) outperformed standalone LLMs (35.4% R@1).
- Performance varied by benchmark composition; higher R@1 with fewer ultra-rare diseases.
Conclusions:
- LLM diagnostic performance for rare diseases is inconsistent and influenced by benchmark characteristics.
- Augmented LLMs and simpler disease profiles improve accuracy.
- High risk of bias and lack of prospective validation necessitate caution before clinical deployment.
Related Concept Videos
Improving Translational Accuracy
3.8K
3.8K
Improving Translational Accuracy
15.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.5K
Sensitivity, Specificity, and Predicted Value
1.8K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.8K

