Related Experiment Video
Updated: Sep 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing Artificial Intelligence Language Models for Patient-Oriented Information on Chiari Malformation Types: A
Mert Çetin1, Ali Çağlar Turgut2, Orhan Beger3
1Class VI student, Gaziantep University Faculty of Medicine, Gaziantep, Türkiye.
Objective:
To evaluate the performance of contemporary large language model (LLM)-based artificial intelligence (AI) systems in providing patient-oriented information across multiple Chiari malformation (CM) subtypes.
Methods:
Five AI models (ChatGPT-4o, Gemini, Copilot, DeepSeek, and Perplexity AI) were evaluated using seven patient-oriented questions on the definition, epidemiology, etiology, symptomatology, diagnosis, treatment, and prognosis of CM. Nine CM subtypes were included: Types 0, 0.5, 1, 1.5, 2, 3, 3.5, 4, and 5. A total of 315 AI-generated responses were independently assessed by three blinded neurosurgeons using a three-point ordinal scoring system for accuracy, comprehensiveness, and conciseness.
Results:
There were significant differences among the AI models across all domains evaluated (all p<0.001). Gemini performed most strongly in accuracy and comprehensiveness, while Perplexity performed best in conciseness. Copilot generally performed worse across the domains evaluated. Significant subtype-based differences were also identified, with AI-generated responses concerning Types 1 and 2 performing substantially better than those concerning rarer subtypes. Diagnostic questions achieved the highest overall performance, while prognosis- and definition-related questions performed comparatively less well. Reviewer-based analyses revealed substantial variability in scoring behavior.
Conclusions:
Contemporary LLMs perform variably in providing patient-oriented information about CMs. Although some AI systems generated relatively accurate and comprehensive responses for well-established CM subtypes, performance was lower for rarer and less clearly defined variants, highlighting the continuing importance of expert oversight in delivering AI-assisted medical information.