Related Experiment Video
Updated: Dec 29, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Combining deep learning with token selection for patient phenotyping from electronic health records
Zhen Yang1, Matthias Dehmer2,3,4, Olli Yli-Harja5,6,7
1Predictive Society and Data Analytics Lab, Tampere University, Tampere, Korkeakoulunkatu 10, 33720, Tampere, Finland.
Deep learning and natural language processing analyze electronic health record (EHR) notes for patient phenotyping. A combined word- and sentence-level CNN improved chronic pain disorder classification.
Area of Science:
- Computational linguistics
- Medical informatics
- Artificial intelligence in healthcare
Background:
- Electronic health records (EHRs) contain vast clinical data, but free-form text notes pose analysis challenges.
- Variability in writing styles across EHRs complicates automated information extraction.
- Patient phenotyping from clinical notes requires advanced analytical methods.
Purpose of the Study:
- To investigate deep learning neural networks combined with natural language processing for analyzing clinical discharge summaries.
- To evaluate the impact of network architectures, sample sizes, and token information content on patient phenotyping accuracy.
- To identify optimal methods for classifying ten patient disorders, including challenging cases like Chronic Pain.
Main Methods:
- Utilized deep learning neural networks and natural language processing (NLP) to analyze text data from clinical discharge summaries.
- Investigated various network architectures, including convolutional neural networks (CNNs) at word-level and combined word- and sentence-level (ws-CNN).
- Analyzed the influence of sample size and token information content, employing a token selection mechanism based on Zipf's law.
Main Results:
- The combined word- and sentence-level input CNN (ws-CNN) demonstrated the largest performance gain for classifying Chronic Pain patients.
- Both data quality and quantity are crucial for leveraging complex network architectures beyond simple word-level CNNs.
- Larger sample sizes are necessary for complex models due to limited information per sample, often carried by few tokens.
Conclusions:
- Deep learning models, particularly ws-CNNs, show significant potential for accurate patient phenotyping from EHR discharge summaries.
- The findings highlight the importance of data quantity and quality in enhancing deep learning model performance for clinical text analysis.
- A token selection mechanism, informed by Zipf's law, aids in understanding token information content and contributes to explainable AI.
More Related Videos
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018