Related Experiment Video
Updated: Jan 29, 2026

Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
Published on: July 2, 2013
AI-driven integration of imaging and radiology language improves stroke mortality prediction
Albara Alotaibi1, Abdalla Moustafa2, Laith Abualloush1
1Department of Radiology, Jordan University of Science and Technology, Irbid, Jordan.
Background:
Quantifying radiologic phenotypes from unstructured text and imaging data may enhance clinical prediction in acute stroke but remains underexplored. This study evaluates the feasibility of automated stroke phenotyping across complementary data sources and assesses whether NLP-derived features from radiology reports improve mortality prediction beyond structured electronic health record (EHR) data.
Methods:
Two complementary datasets were analyzed. First, MRI lesion masks from a public dataset (n = 60) were processed using NiBabel to calculate lesion volumes and characterize imaging features as a quantitative reference cohort. Second, 15,492 head CT/MRI reports from the MIMIC-III database were processed through a rule-based NLP pipeline to identify eight key stroke phenotypes: hemorrhage, infarct, midline shift, edema, chronic change, and vascular territories (ACA, MCA, PCA). Each classifier was trained on TF-IDF features using logistic regression and evaluated by AUC and F1-score. Probabilistic outputs from the NLP models were merged with structured admission data (age, sex, ICD-9 codes, length of stay) to predict in-hospital mortality in 3,999 stroke admissions using logistic regression.
Results:
In the MRI reference cohort, mean lesion volume was 3.85 × 106 mm3 ± 4.83 × 105, demonstrating the feasibility of automated lesion quantification from open datasets. Across the MIMIC-III cohort, the best NLP models achieved AUC = 0.974 (hemorrhage) and 0.957 (edema) with balanced F1-scores (0.945 and 0.891, respectively). Incorporating text-derived phenotypes into mortality models improved discrimination modestly (AUC 0.616 vs. 0.558; ΔAUC = +0.058). Permutation analysis revealed ICD-9 codes (ΔAUC = 0.091), edema (0.051), and infarct (0.046) as top contributors to mortality prediction.
Conclusion:
Automated extraction of stroke phenotypes from both quantitative imaging and radiology text is feasible and reproducible across open datasets. Although MRI lesion volume was not incorporated into mortality models due to dataset limitations, NLP-derived radiologic phenotypes from clinical text provided complementary, interpretable information beyond structured EHR data and modestly improved mortality risk stratification. These findings support the potential of text-derived imaging phenotypes as a scalable and clinically practical enhancement to stroke outcome prediction. Broader validation, incorporation of additional modalities including direct imaging metrics, and workflow-aware implementation strategies will be important next steps toward translating these models into actionable clinical decision support.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Predicting Molecular Geometry
Radiological Investigation I: X-ray and CT