Related Experiment Video
Updated: Oct 25, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Developing a Natural Language Processing tool to identify perinatal self-harm in electronic healthcare records
Karyn Ayre1,2, André Bittar3, Joyce Kam4
1Section of Women's Mental Health, Health Service and Population Research Department, Institute of Psychiatry, Psychology and Neuroscience, Kings College London, London, United Kingdom.
This study developed a Natural Language Processing (NLP) tool to accurately identify perinatal self-harm in electronic healthcare records (EHRs). The tool shows promise for improving the detection of self-harm during pregnancy and the postpartum period.
Area of Science:
- Computational Linguistics
- Clinical Informatics
- Public Health
Background:
- Perinatal self-harm is under-researched, with current prevalence estimates likely underestimated.
- Electronic healthcare records (EHRs) offer a valuable data source for studying perinatal self-harm.
- Methodological limitations hinder accurate identification of self-harm in this population.
Purpose of the Study:
- To develop a Natural Language Processing (NLP) tool for identifying perinatal self-harm mentions in EHRs with high precision and recall.
- To utilize the NLP tool to identify individuals who have experienced perinatal self-harm through their EHR data.
Main Methods:
- Extracted de-identified EHRs from secondary mental healthcare service users.
- Developed an NLP tool using spaCy for linguistic processing of EHR data.
- Evaluated mention-level performance (span, status, temporality, polarity) against a manual coding standard.
- Assessed service-user level performance, including the impact of a heuristic rule.
Main Results:
- Mention-level performance achieved F-scores, precision, and recall above 0.8 for span, polarity, and temporality.
- Status and temporality showed lower agreement (Cohen's kappa 0.68 and 0.62, respectively).
- Service-user level performance with a heuristic rule yielded a macro-averaged F-score of 0.81, with a positive likelihood ratio of 9.4.
Conclusions:
- Developing an NLP tool to identify perinatal self-harm in EHRs is feasible and demonstrates acceptable validity.
- The tool has limitations concerning the temporality of self-harm events.
- The NLP tool, enhanced with a heuristic rule, can effectively identify individuals at the service-user level.
Related Concept Videos
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Methods of Documentation VII: EMR

