Related Experiment Video
Updated: Apr 27, 2026

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
Addressing diagnostic code variability in intimate partner violence surveillance through natural language processing:
Daniel R Harris1, Benjamin P Clements2, Nicholas Anthony1
1Institute for Pharmaceutical Outcomes & Policy, Department of Pharmacy Practice and Science, College of Pharmacy, University of Kentucky, Lexington, KY 40508, USA; Institute for Biomedical Informatics, University of Kentucky, Lexington, KY 40508, USA.
Background:
Intimate partner violence (IPV) and substance use disorders (SUDs) represent major public health challenges, yet identifying IPV cases in electronic health records (EHR) remains difficult due to inconsistent documentation practices and variable billing procedures.
Methods:
We constructed three patient cohorts based on SUDs or overdoses involving stimulants, opioids, or both from the EHR. Using the Open Health Natural Language Processing (OHNLP) toolkit, we developed a rule-based classifier to identify IPV in clinical notes. We validated classifier performance through manual chart review and linked cohorts to Kentucky's fatal overdose surveillance system to examine the relationship between IPV and fatal overdose risk.
Results:
We analyzed 15,557,678 clinical notes from 29,447 patients with SUDs (2017-2023). Our classifier demonstrated strong performance: recall 0.91, precision 0.85, F1-score 0.89. IPV was detected in 5.6% of patients via NLP compared to 0.2-2.5% using various ICD-10-CM diagnostic code definitions; this variation reflected different levels of code specificity across published definitions. Notably, 90% of NLP-detected cases lacked corresponding diagnostic codes. The cohort with both stimulant and opioid disorders showed the highest IPV prevalence (7.5%), followed by opioid-only (5.2%) and stimulant-only (4.8%). Among fatal overdose cases, IPV documentation rates (4.4% via NLP, 5.1% via diagnostic codes) were similar to the overall cohort, suggesting missed intervention opportunities.
Conclusion:
NLP-based analysis of EHR clinical notes identified substantially more IPV cases than diagnostic codes alone. The elevated IPV prevalence among people with polysubstance use highlights a particularly vulnerable population requiring integrated screening and intervention strategies. Routine NLP surveillance could significantly improve IPV case identification in populations with SUDs.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
04:04Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Related Concept Videos
Diagnostic and Statistical Manual of Mental Disorders (DSM)
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
Sex Linked Disorders
Automatic Processing and Automatic Social Behavior