Related Experiment Video
Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparison of ChatGPT and DeepSeek large language models in the diagnosis of pericarditis
Aman Goyal1, Samia Aziz Sulaiman2, Abdallah Alaarag2
1Department of Internal Medicine, Cleveland Clinic Foundation, Cleveland, OH 44195, United States.
Background:
The integration of sophisticated large language models (LLMs) into healthcare has recently garnered significant attention due to their ability to leverage deep learning techniques to process vast datasets and generate contextually accurate, human-like responses. These models have been previously applied in medical diagnostics, such as in the evaluation of oral lesions. Given the high rate of missed diagnoses in pericarditis, LLMs may support clinicians in generating differential diagnoses-particularly in atypical cases where risk stratification and early identification are critical to preventing serious complications such as constrictive pericarditis and pericardial tamponade.
Aim:
To compare the accuracy of LLMs in assisting the diagnosis of pericarditis as risk stratification tools.
Methods:
A PubMed search was conducted using the keyword "pericarditis", applying filters for "case reports". Data from relevant cases were extracted. Inclusion criteria consisted of English-language reports involving patients aged 18 years or older with a confirmed diagnosis of acute pericarditis. The diagnostic capabilities of ChatGPT o1 and DeepThink-R1 were assessed by evaluating whether pericarditis was included in the top three differential diagnoses and as the sole provisional diagnosis. Each case was classified as either "yes" or "no" for inclusion.
Results:
From the initial search, 220 studies were identified, of which 16 case reports met the inclusion criteria. In assessing risk stratification for acute pericarditis, ChatGPT o1 correctly identified the condition in 10 of 16 cases (62.5%) in the differential diagnosis and in 8 of 16 cases (50.0%) as the provisional diagnosis. DeepThink-R1 identified it in 8 of 16 cases (50.0%) and 6 of 16 cases (37.5%), respectively. ChatGPT o1 demonstrated higher accuracy than DeepThink-R1 in identifying pericarditis.
Conclusion:
Further research with larger sample sizes and optimized prompt engineering is warranted to improve diagnostic accuracy, particularly in atypical presentations.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
08:56Sterile Pericarditis in Aachener Minipigs As a Model for Atrial Myopathy and Atrial Fibrillation
Published on: September 24, 2021
Related Concept Videos
Pericarditis II: Clinical Features and Diagnostic Tests
Pericarditis I: Introduction
Rheumatic Heart Disease II: Clinical Manifestations and Diagnostic Studies
Myocarditis II: Clinical Features and Diagnostic Tests
Pericarditis III: Medical Management
Myasthenia Gravis: Diagnostic Tests
The edrophonium test is a diagnostic tool for myasthenia gravis. It involves...