Related Experiment Video
Updated: Jan 29, 2026

Enumeration of Major Peripheral Blood Leukocyte Populations for Multicenter Clinical Trials Using a Whole Blood Phenotyping Assay
Published on: September 16, 2012
Comparative Performance of Multimodal and Unimodal Large Language Models Versus Multicenter Human Clinical Experts in
1Department of Emergency Medicine, Etimesgut Sehit Sait Erturk State Hospital, Ankara 06790, Turkey.
Abstract:
Background: Multimodal large language models (MLLMs) integrating multiple AI systems and unimodal large language models (LLMs) represent distinct approaches to clinical decision support. Their comparative performance against human clinical experts in complex cardiovascular emergencies remains inadequately characterized. Objective: To compare the performance of a combined MLLM system (GPT-4V + Med-PaLM 2 + BioGPT), a unimodal LLM (ChatGPT-5.2), and human physicians from multiple centers (radiologists, emergency medicine specialists, cardiovascular surgeons) on aortic dissection clinical questions across diagnosis, treatment, and complication management domains. Methods: This multicenter cross-sectional study was conducted across five tertiary care centers in Turkey (Elazığ, Ankara, Antalya). A total of 25 validated multiple-choice questions were categorized into three domains: diagnosis (n = 8), treatment (n = 9), and complication management (n = 8). Questions were administered to the MLLM, ChatGPT-5.2 (Unimodal), and nine physicians from five centers: radiologists (n = 3), emergency medicine specialists (n = 3), and cardiovascular surgeons (n = 3). Statistical comparisons utilized chi-square tests. Results: Overall accuracy was 92.0% for the MLLM and 96.0% for ChatGPT-5.2 (Unimodal). Among human physicians, cardiovascular surgeons achieved 96.0%, radiologists 92.0%, and emergency medicine specialists 89.3%. The MLLM excelled in diagnosis (100%) but showed lower performance in treatment (88.9%) and complication management (87.5%). No significant differences were observed between AI models and human physician groups (all p > 0.05). Conclusions: Both the MLLM and unimodal ChatGPT-5.2 demonstrated performance within the range of human clinical experts in this controlled assessment of aortic dissection scenarios, though definitive conclusions regarding equivalence require larger-scale validation. These findings support further investigation of complementary roles for different AI architectures in clinical decision support.
Related Concept Videos
Aortic Regurgitation III: Medical Management
Aortic Regurgitation IV: Nursing Management
Esophageal Perforation-II: Clinical Manifestations and Management
Clinical Manifestations:
Barrett Esophagus-II: Clinical Manifestations and Management
To diagnose Barrett's esophagus, healthcare providers often recommend an endoscopy for those showing symptoms of acid reflux. The procedure...
Esophageal Varices-II: Clinical Features and Management
In the initial assessment, a thorough review of the patient's medical history is vital to identify risk factors such as liver disease, alcohol...
Gastritis III: Clinical Manifestations and Management
Clinical manifestations of acute gastritis
The patient with acute gastritis may have a rapid onset of symptoms, such as epigastric pain or discomfort, dyspepsia, anorexia, hiccups, or nausea and vomiting, which can last from a few hours to a few days. Erosive or hemorrhagic gastritis may cause bleeding, which may manifest as blood in vomit or as...

