Related Experiment Video
Updated: May 20, 2026

Surgical Training for the Implantation of Neocortical Microelectrode Arrays Using a Formaldehyde-fixed Human Cadaver Model
Published on: November 19, 2017
CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications
Anton Alyakin1,2,3, Jaden Stryker1, Daniel Alexander Alber1,4
1Department of Neurological Surgery, NYU Langone Health, New York, New York, USA.
Background And Objectives:
General purpose vision-language models (VLMs) demonstrate impressive capabilities, but their opaque training on uncurated internet data poses critical limitations for high-stakes decision making, such as in neurosurgery. We present CNS-Obsidian, a neurosurgical VLM trained on peer-reviewed neurosurgical literature, and demonstrate its clinical utility compared with GPT-4o in a real-world setting.
Methods:
We compiled 23 984 articles from Neurosurgery Publications journals, yielding 78 853 figures and captions. Using GPT-4o and Claude Sonnet-3.5, we converted these image-text pairs into 263 064 training samples across 3 formats: instruction fine-tuning, multiple-choice questions, and differential diagnosis. We trained CNS-Obsidian, a fine-tune of the 34-billion parameter Large Language and Visual Assistant-Next model. In a blinded, randomized deployment trial at NYU Langone Health (August 30-November 30, 2024), neurosurgeons were assigned to use either CNS-Obsidian or a Health Insurance Portability and Accountability Act-compliant GPT-4o end point as a diagnostic copilot after patient consultations. Primary outcomes were diagnostic helpfulness and accuracy, assessed through user ratings and presence of the correct diagnosis within the VLM-provided differential, respectively.
Results:
CNS-Obsidian matched GPT-4o on synthetic questions (76.13% vs 77.54%, P = .235), but only achieved 46.81% accuracy on human-generated questions vs GPT-4o's 65.70% (P < 10-15). In the randomized trial, 70 consultations were evaluated (32 CNS-Obsidian, 38 GPT-4o) from 959 total consults (7.3% utilization). CNS-Obsidian received positive ratings in 40.62% of cases vs 57.89% for GPT-4o (P = .230). Both models included correct diagnosis in approximately 60% of cases (59.38% vs 65.79%, P = .626).
Conclusion:
Domain-specific VLMs trained on curated scientific literature can approach frontier model performance in specialized medical domains despite being orders of magnitude smaller and less expensive to train. This establishes a transparent framework for scientific communities to build specialized artificial intelligence models. However, low clinical utilization suggests chatbot interfaces may not align with specialist workflows, indicating need for alternative artificial intelligence integration strategies.
More Related Videos
Related Concept Videos
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Neurons as Communicators of the Brain
Cell Body
The cell body, also known...
Organization of the Brain
Hindbrain
The hindbrain, located at the base of the brain, plays a vital role in regulating automatic processes that sustain life. It includes the medulla oblongata, which is essential for...
Cerebrospinal Fluid
CSF Production
CSF is produced mainly in the choroid plexus, a network of capillaries and ependymal cells located within the ventricular system of the brain.
Neuronal Communication

