Related Experiment Video
Updated: Jun 26, 2026

06:18
Pioneering Patient-Specific Approaches for Precision Surgery Using Imaging and Virtual Reality
Published on: April 5, 2024
Evaluating expert decision-pattern alignment in endovascular planning for intracranial aneurysms using a multimodal
Mustafa Demir1, Yunus Yasar2, Yusuf Agackaya3
1Department of Interventional Radiology, Ümraniye Eğitim ve Araştırma Hastanesi, Ümraniye, Turkey. drmstfdmr1@gmail.com.
Neuroradiology
|June 25, 2026
Summary
A multimodal large language model (LLM) showed moderate agreement with expert decisions for endovascular treatment of intracranial aneurysms. The LLM reproduced expert heuristics, suggesting potential as a tool for standardizing treatment strategies.
Area of Science:
- Neurosurgery
- Medical Imaging
- Artificial Intelligence
Background:
- Intracranial aneurysm treatment planning involves complex factors like vascular geometry and patient risk.
- Current treatment selection relies on expert heuristics due to multiple viable strategies.
- The ability of multimodal large language models (LLMs) to replicate these expert decisions is not well-understood.
Purpose of the Study:
- To evaluate if a multimodal LLM can reproduce expert decision-making patterns in endovascular treatment planning for intracranial aneurysms.
- To compare the LLM's performance against an independent expert and a neurointerventional trainee.
Main Methods:
- A retrospective study analyzed 59 patients with unruptured intracranial aneurysms treated with endovascular therapy.
- A multimodal LLM (GPT-5), an independent neuroradiologist, and a neurointerventional fellow evaluated cases against a neurovascular board's decisions.
- Evaluators received standardized clinical summaries and 3D-DSA projections; agreement was quantified using Cohen's kappa.
Main Results:
- GPT-5 demonstrated moderate concordance (κ=0.64) with board-selected treatment modalities, within expert variability.
- The LLM's agreement exceeded that of the trainee (κ=0.15).
- GPT-5 proposals were often preferred or acceptable alternatives, particularly in strategic planning, device sizing, and landing-zone estimation.
Conclusions:
- A multimodal LLM can generate endovascular treatment strategies aligning with expert heuristics in a clinical context.
- The model appears to replicate codified expert knowledge rather than performing autonomous reasoning.
- LLMs show promise as adjunct tools for standardizing neurointerventional decision-making and supporting trainee education, pending further validation.