Related Experiment Video
Updated: Jan 2, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Identifying Transitional High Cost Users from Unstructured Patient Profiles Written by Primary Care Physicians
Haoran Zhang1,2,3, Elisa Candido, Andrew S Wilton
1Department of Computer Science, University of Toronto, Canada.
Predicting patients at risk of becoming High Cost Users (HCUs) can save healthcare costs. Using specialized word embeddings from electronic health records improved prediction accuracy, especially for identifying new HCUs.
Area of Science:
- Health Informatics
- Machine Learning in Healthcare
- Clinical Data Analysis
Background:
- Identifying High Cost Users (HCUs) enables proactive interventions, improving patient outcomes and reducing healthcare expenditures.
- Electronic medical records contain valuable unstructured text data crucial for predicting patient risk.
- Traditional methods often struggle with the nuances of clinical notes, including domain-specific language.
Purpose of the Study:
- To predict the 2016 High Cost User (HCU) status of patients using 2015 electronic medical record data.
- To evaluate the effectiveness of context-specific word embeddings versus pre-trained embeddings for this prediction task.
- To assess the performance improvement of an aggregated word embedding model (EmbEncode) over traditional bag-of-words approaches.
Main Methods:
- Utilized free-form text data from 2015 cumulative patient profiles in Ontario family care practices.
- Developed context-specific word embeddings from unstructured clinical notes, accounting for domain-specific abbreviations and spellings.
- Compared a novel EmbEncode model against a bag-of-words model using held-out AUROC metrics.
Main Results:
- Context-specific word embeddings significantly outperformed pre-trained embeddings (Wikipedia, MIMIC, Pubmed).
- The EmbEncode model achieved a higher held-out AUROC (82.48±0.35%) compared to bag-of-words (81.85±0.36%, p = 3.2 × 10-4).
- EmbEncode used substantially fewer features and non-zero coefficients, indicating greater efficiency.
- Removing recurrent HCUs from the training set enhanced the prediction of new HCUs with minimal impact on predicting recurrent ones.
Conclusions:
- Aggregated word embeddings derived from clinical text (EmbEncode) offer a more effective and efficient method for predicting High Cost Users.
- Tailoring feature extraction to the specific clinical context is crucial for improving predictive model performance.
- A refined approach focusing on transitional HCUs can optimize intervention strategies within healthcare systems.
More Related Videos
Related Concept Videos
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Analysis of Population Pharmacokinetic Data
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Secondary Healthcare System

