Related Experiment Video
Updated: Apr 4, 2026

06:11
Arterial Pouch Microsurgical Bifurcation Aneurysm Model in the Rabbit
Published on: May 14, 2020
2.9K
Benchmarking LLM decision support in inflammatory aneurysms: DeepSeek-R1, DeepSeek-V3 and ChatGPT-4o
Jianqiang Hao1, Changwei Zhang2
1Department of Neurosurgery, West China Hospital, Sichuan University, Chengdu, China.
Clinical and Experimental Medicine
|April 3, 2026
Summary
Large language models (LLMs) show promise for inflammatory aneurysm decisions. DeepSeek-R1 and ChatGPT-4o offer high accuracy, while DeepSeek-V3 is faster but less accurate, highlighting the need to match model choice to clinical criticality.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Vascular Surgery Decision Support
Background:
- Guideline-concordant decision support is crucial for managing inflammatory aneurysms.
- Evaluating the accuracy and efficiency of large language models (LLMs) in this critical domain is necessary.
Purpose of the Study:
- To compare the diagnostic, interventional, and surveillance decision-making accuracy and time efficiency of three LLMs: DeepSeek-R1, ChatGPT-4o, and DeepSeek-V3.
- To assess LLM performance against contemporary aortic guidance.
Main Methods:
- A prospective, rater-blinded evaluation using a 50-item instrument aligned with aortic guidelines.
- Five vascular specialists scored LLM outputs using a 0/1/2 rubric.
- Analysis included total accuracy, completion time, and domain-specific subscores, with inter-rater reliability assessed.
Main Results:
- Mean total accuracy scores were 89.8% (DeepSeek-R1), 88.4% (ChatGPT-4o), and 77.8% (DeepSeek-V3), with R1 and 4o showing near-parity and significantly outperforming V3.
- Completion times varied significantly: DeepSeek-V3 (19.8s), ChatGPT-4o (33.4s), and DeepSeek-R1 (58.5s).
- Accuracy correlated positively with completion time (Pearson r=0.802), with differences concentrated in guideline-derived items.
Conclusions:
- In complex, safety-critical decisions for inflammatory aneurysms, slower, reasoning-intensive LLMs achieve higher accuracy.
- A low-latency multimodal system (ChatGPT-4o) provides near-parity accuracy at significantly reduced times.
- Model selection should align with task criticality, and human oversight remains essential.
Related Concept Videos
Aneurysm III: Interprofessional Care
450
Aneurysm management involves either conservative medical therapy or surgical intervention, depending on the size and symptoms of the aneurysm. Conservative management is generally reserved for smaller, asymptomatic aneurysms, while larger or symptomatic aneurysms often necessitate surgical repair.Conservative Medical TherapyFor small, asymptomatic aneurysms, particularly abdominal aortic aneurysms (AAA) less than 5.5 centimeters in diameter, conservative medical therapy is recommended. This...
450
Aneurysm II: Clinical Manifestations and Diagnostic Studies
526
Thoracic, aortic arch and abdominal aneurysms are significant vascular conditions that can present with various clinical manifestations and lead to serious complications. Understanding these manifestations and the appropriate diagnostic studies is essential for effective management and treatment.Thoracic Aortic AneurysmsThoracic aortic aneurysms often remain asymptomatic until they reach a size that impinges on adjacent structures. They typically cause deep, diffuse chest pain that radiates to...
526
Aneurysm IV: Nursing Management
582
Vigilant monitoring for aneurysm rupture is essential for patients undergoing aortic surgery.Preoperative Nursing ManagementContinuously monitor the patient for manifestations of aneurysm rupture, such as pallor, weakness, tachycardia, hypotension, abdominal, back, groin, or periumbilical pain, changes in consciousness, and a pulsating abdominal mass. Regularly assess the patient's peripheral pulses.Instruct the patient to consume a clear liquid diet the day before surgery and administer...
582

