Related Experiment Video
Updated: Sep 19, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of ChatGPT-4o and Four Open-Source Large Language Models in Generating Diagnoses Based on China's Rare
Wei Zhong1, YiFan Liu1, Yan Liu1
1Department of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, No. 251 Yaojiayuan Road, Chaoyang District, Beijing, China, 8618810963279.
ChatGPT-4o achieved the highest diagnostic accuracy for rare diseases. Integrating retrieval augmented generation (RAG) significantly improved open-source large language models (LLMs), highlighting the need for careful model selection based on language and parameterization for clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Rare Disease Diagnostics
Background:
- Diagnosing rare diseases is complex, often limited by physician knowledge.
- Large language models (LLMs) present a novel approach to enhance diagnostic workflows.
Purpose of the Study:
- Evaluate diagnostic accuracy of ChatGPT-4o and open-source LLMs for rare diseases.
- Assess the impact of language (English vs. Chinese) on LLM diagnostic performance.
- Explore the effectiveness of retrieval augmented generation (RAG) and chain-of-thought (CoT) reasoning.
Main Methods:
- Extracted clinical manifestations for 121 rare diseases from China's rare disease catalog.
- Assessed ChatGPT-4o and four open-source LLMs (qwen2.5:7b, Llama3.1:8b, qwen2.5:72b, Llama3.1:70b) in English and Chinese.
- Re-evaluated the lowest-performing model using RAG and CoT; compared accuracy using McNemar test; surveyed clinicians on rare disease familiarity.
Main Results:
- ChatGPT-4o achieved the highest diagnostic accuracy (90.1%).
- Language significantly impacted LLM performance; Llama3.1:8b performed better in English, while qwen2.5:7b showed similar performance in both languages.
- RAG improved qwen2.5:7b accuracy to 79.3% with 85.9% retrieval precision; clinician surveys revealed knowledge gaps in rare disease management.
Conclusions:
- ChatGPT-4o demonstrated superior diagnostic performance for rare diseases.
- Open-source LLMs require larger parameters for comparable accuracy in Chinese; RAG integration enhances diagnostic utility, especially in resource-constrained settings.
- Prioritize language relevance and curated knowledge bases for clinical LLM implementation, exercising caution with low-parameter models.
More Related Videos
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Single Nucleotide Polymorphisms-SNPs

