Related Experiment Video
Updated: Mar 10, 2026

Author Spotlight: Bridging Gaps in Anatomy and Establishing a Foundation for Algorithmic Studies
Published on: December 15, 2023
Large-scale identification of patients with cerebral aneurysms using natural language processing
Victor M Castro1, Dmitriy Dligach1, Sean Finan1
1From Research Information Systems and Computing (V.M.C., V.G., S.M.), Partners Healthcare; Boston Children's Hospital Informatics Program (D.D., S.F., G.S.); Harvard Medical School (D.D., S.Y., A.C., M.A.-E.-B., N.A.S., S.M., S.T.W., R.D.); Department of Medicine (S.Y., S.T.W.), Department of Neurosurgery (A.C., M.A.-E.-B., R.D.), Division of Rheumatology, Immunology and Allergy (N.A.S.), and Channing Division of Network Medicine (S.T.W., R.D.), Brigham and Women's Hospital, Boston, MA; Center for Statistical Science (S.Y.), Tsinghua University, Beijing, China; Department of Neurology (S.M.), Massachusetts General Hospital; and Biostatistics (T.C.), Harvard School of Public Health, Boston, MA.
Natural language processing (NLP) applied to electronic medical records (EMR) accurately identified patients with cerebral aneurysms. This method efficiently created a large cohort for research, enabling new studies on brain aneurysms.
Area of Science:
- Biomedical Informatics
- Medical Natural Language Processing
- Clinical Data Mining
Background:
- Electronic medical records (EMR) contain vast patient data but require sophisticated tools for accurate disease identification.
- Identifying specific patient cohorts, such as those with cerebral aneurysms, is crucial for epidemiological and genetic research.
- Traditional methods of patient identification can be time-consuming and prone to errors.
Purpose of the Study:
- To develop and validate a natural language processing (NLP) algorithm for precise identification of cerebral aneurysm patients within an EMR system.
- To create a matched cohort of patients with cerebral aneurysms and control subjects using EMR data.
- To assess the accuracy and efficiency of the NLP-driven approach compared to traditional coding methods.
Main Methods:
- Utilized ICD-9 and Current Procedural Terminology codes to extract an initial patient dataset from the EMR.
- Trained a classification algorithm using NLP techniques, employing .632 bootstrap cross-validation to mitigate overfitting bias.
- Applied the validated algorithm to the full dataset to classify patients and matched controls based on age, sex, race, and healthcare utilization.
Main Results:
- Identified 55,675 patients with initial codes suggestive of cerebral aneurysms from over 4.2 million patient records.
- An NLP algorithm, incorporating 8 coded and 14 NLP variables, achieved an area under the receiver-operating characteristic curve of 0.95.
- The final classification identified 5,589 patients with cerebral aneurysms and 54,952 matched controls, with a positive predictive value of 0.86 in a validation cohort.
Conclusions:
- Leveraged EMR data with NLP to successfully establish a large, well-defined cohort of patients with intracranial aneurysms and their controls.
- The developed NLP algorithm demonstrates high accuracy and efficiency in patient cohort identification.
- This methodology is generalizable and can be applied to identify patient cohorts for various diseases, facilitating broader epidemiological and genetic research.

