Related Experiment Video
Updated: Jan 30, 2026

Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
Natural Language Processing to Automate Cerebrovascular Event Identification in Stroke Alerts
Asala N Erekat1,2, Laura K Stein2, Bradley N Delman3
1Clinical Neuro-Informatics Center Icahn School of Medicine at Mount Sinai New York NY.
Background:
Frequent false-positive stroke alerts can strain health resources. Machine learning models can predict stroke alert accuracy and potentially reduce this strain, but these models require time-consuming labeling of large data sets. Weak labeling can accelerate machine learning model development by assigning annotations based on expert-defined heuristic rules rather than manual review. We sought to label a large, unlabeled sample of stroke alerts according to a binary outcome (presence/absence of acute cerebrovascular disease) using weak labeling.
Methods:
We developed a weak labeling heuristic ensemble consisting of 4 hierarchical tiers, each of which generated a binary label using custom labeling algorithms. Tier 1 used rule-based named-entity recognition to generate binary labels from brain radiology reports. Tier 2 aggregated Tier 1 outputs over a 48-hour window. Tier 3 determined labels using diagnosis codes from stroke alert hospital encounters. Tier 4 generated a "final" encounter-level label based on label output combinations of Tiers 2 and 3. In 3 separate samples of stroke alerts, we determined sensitivity, specificity, and F1 scores of Tiers 1, 2, 3, and 4 by comparing Tier outputs to manual chart review.
Results:
We identified 16 512 stroke alert activations between 2011 and 2021. For Tier 1, performance metrics were based on an initial manual review of 300 neuroimaging reports, achieving a sensitivity of 0.84, specificity of 0.96, and an F1 of 0.87. Tier 2 incorporated 716 neuroimaging reports with a sensitivity of 0.93, specificity of 0.90, and an F1 of 0.91. Tiers 3 and 4 were validated against 250 encounters. Tier 3 achieved a sensitivity of 0.77, specificity of 0.89, and an F1 of 0.80. Tier 4 achieved a sensitivity of 0.92, specificity of 0.86, and an F1 of 0.87.
Conclusions:
We successfully labeled a large registry of stroke alerts using weak labeling. This framework can potentially be extended to other clinical data sets.
More Related Videos
07:30Evaluation of the Cognitive Performance of Hypertensive Patients with Silent Cerebrovascular Lesions
Published on: April 23, 2021
10:15Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
Published on: July 2, 2013
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
What is Natural Selection?
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Nature and Nurture