Related Experiment Video
Updated: Jan 16, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Testing regular expression searches and machine learning models to determine housing instability and low income
Stephanie Garies1,2,3, Christopher Meaney4, Karen Weyman5,4
1Department of Family Medicine, University of Calgary, Calgary, AB, Canada. sgaries@ucalgary.ca.
Background:
Housing and income are important social determinants of health (SDoH). Primary care providers often do not have information about these determinants, which could be used to support equitable health system planning and care delivery. The aim of this study was to use primary care electronic medical record (EMR) data to test two approaches (machine learning and regular expression searches) to obtain information about patients' housing instability and low income status.
Methods:
We used de-identified EMR data from the St. Michael's Hospital Academic Family Health Team (Toronto, Ontario, Canada). A Health Equity Questionnaire is also routinely distributed to patients and includes questions about income and housing status; this formed the reference standard. First, a regular expression (REGEX) classifier was created using key text terms and codes; the second approach used supervised machine learning models (XGBoost). Discrimination and calibration metrics were calculated as compared to the patient-reported responses.
Results:
11,794 eligible patients were included in the housing cohort and 10,454 were in the income cohort. Overall, both approaches had poor sensitivity for determining both housing instability (XGBoost: 3.1%, REGEX: 29.0%) and low income status (XGBoost: 41.7%, REGEX: 17.6%). Positive predictive value (PPV) was satisfactory for the machine learning approach (83.3% for housing, 72.9% for income).
Conclusion:
While the machine learning approach demonstrated reasonable PPV, the overall metrics were poor and unlikely to be useful in a clinical setting for identifying patients with housing or economic needs. More robust analysis could be explored, but continued patient-captured SDoH information is necessary.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Steps in Outbreak Investigation
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Statistical Software for Data Analysis and Clinical Trials

