Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study
Jie Song1, Zhichuan Xu1, Meng Xiao1
1Department of Ophthalmology and Institutes for Systems Genetics, Frontiers Science Center for Disease-related Molecular Network, West China Hospital of Sichuan University, No.2222, Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610041, China, 86 15995854635.
Journal of Medical Internet Research
|July 23, 2026
Summary
Large language models (LLMs) show promise for rare disease diagnosis. They achieved higher accuracy than humans in initial specialty triage, offering faster and more consistent results for complex cases.
Area of Science:
- Medical Diagnostics
- Artificial Intelligence in Healthcare
- Rare Disease Research
Background:
- Specialty triage is crucial for rare disease diagnosis but often challenging due to complex patient presentations.
- Overlapping, multisystem, and atypical symptoms complicate initial specialty selection, potentially delaying diagnosis.
Purpose of the Study:
- To assess the accuracy, response time, and consistency of large language models (LLMs) for rare disease specialty triage.
- To compare LLM performance against registered nurses and nonmedical participants.
Main Methods:
- Retrospective benchmarking using five rare disease datasets.
- Evaluation of fourteen LLMs over five independent runs per case.
- Performance metrics included accuracy, response time, and consistency, with subgroup analyses.
Main Results:
- LLM accuracy ranged from 0.4378 to 0.7141, with Claude-opus-4-5 achieving the highest accuracy (0.7141).
- GPT-5.1 demonstrated the fastest response time (3.39 s/case) with high accuracy (0.6948).
- LLMs outperformed human participants (registered nurses and nonmedical individuals) in accuracy (0.5978 vs. 0.4914/0.4573).
Conclusions:
- LLMs demonstrate potential as assistive tools for initial specialty triage in rare diseases.
- Model choice, reasoning mode, and phenotype information impact LLM performance.
- Further research is needed to evaluate LLM-based triage in clinical settings and develop supervised workflows.
