Related Experiment Video
Updated: Nov 2, 2025

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
Developing automated methods for disease subtyping in UK Biobank: an exemplar study on stroke
Kristiina Rannikmäe1,2, Honghan Wu3,4, Steven Tominey5
1Centre for Medical Informatics, University of Edinburgh, NINE Edinburgh BioQuarter, 9 Little France Road, Edinburgh, EH16 4UX, UK. kristiina.rannikmae@ed.ac.uk.
Insights
Automated analysis of radiology reports accurately subtypes hemorrhagic strokes (ICH and SAH) and can improve health data research. This method enhances phenotyping for better health outcomes.
Area of Science:
- Medical informatics
- Neurology
- Radiology
Background:
- Routinely collected coded health data requires better phenotyping for research and health improvement.
- Current coded data for hemorrhagic stroke (intracerebral hemorrhage [ICH] and subarachnoid hemorrhage [SAH]) has low precision (<50%).
Purpose of the Study:
- To investigate the feasibility and added value of automated methods using clinical radiology reports to improve stroke subtyping.
- To enhance the accuracy of stroke subtyping beyond existing coded data limitations.
Main Methods:
- Utilized natural language processing and clinical knowledge inference on brain scan reports from UK Biobank participants.
- Assigned stroke subtypes (ischemic, ICH, SAH) and assessed performance using precision and recall at entity and patient levels.
Main Results:
- Automated methods achieved high patient-level precision and recall for ICH (89%) and SAH (82%).
- Performance for ischemic stroke was lower (73% precision, 64% recall), indicating coded data may be preferred for this subtype.
- Entity-level precision and recall ranged from 78% to 100%.
Conclusions:
- Automated analysis of radiology reports offers a feasible, scalable, and accurate solution for improving disease subtyping.
- This method, when combined with administrative coded health data, enhances phenotyping for research and health improvement.
- Further validation in diverse populations is recommended.
Background:
Better phenotyping of routinely collected coded data would be useful for research and health improvement. For example, the precision of coded data for hemorrhagic stroke (intracerebral hemorrhage [ICH] and subarachnoid hemorrhage [SAH]) may be as poor as < 50%. This work aimed to investigate the feasibility and added value of automated methods applied to clinical radiology reports to improve stroke subtyping.
Methods:
From a sub-population of 17,249 Scottish UK Biobank participants, we ascertained those with an incident stroke code in hospital, death record or primary care administrative data by September 2015, and ≥ 1 clinical brain scan report. We used a combination of natural language processing and clinical knowledge inference on brain scan reports to assign a stroke subtype (ischemic vs ICH vs SAH) for each participant and assessed performance by precision and recall at entity and patient levels.
Results:
Of 225 participants with an incident stroke code, 207 had a relevant brain scan report and were included in this study. Entity level precision and recall ranged from 78 to 100%. Automated methods showed precision and recall at patient level that were very good for ICH (both 89%), good for SAH (both 82%), but, as expected, lower for ischemic stroke (73%, and 64%, respectively), suggesting coded data remains the preferred method for identifying the latter stroke subtype.
Conclusions:
Our automated method applied to radiology reports provides a feasible, scalable and accurate solution to improve disease subtyping when used in conjunction with administrative coded health data. Future research should validate these findings in a different population setting.

