Related Experiment Video
Updated: Apr 13, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
8.1K
Toward high-throughput phenotyping: unbiased automated feature extraction and selection from knowledge sources.
Sheng Yu1, Katherine P Liao2, Stanley Y Shaw3
1Partners HealthCare Personalized Medicine, Boston, MA, USA Brigham and Women's Hospital, Boston, MA, USA Harvard Medical School, Boston, MA, USA syu7@partners.org.
Summary
Automated feature extraction from electronic health records (EHRs) using natural language processing (NLP) creates accurate patient phenotyping algorithms. This method improves efficiency and rivals expert-curated features for clinical and genetic research.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Clinical Research
Background:
- Electronic health records (EHRs) contain valuable narrative data for population-scale phenotyping.
- Current methods for selecting text features for phenotyping algorithms are time-consuming and require expert input.
- Developing efficient and accurate phenotyping tools is crucial for advancing clinical and genetic research.
Purpose of the Study:
- To introduce an automated method for extracting and selecting informative text features from EHRs for phenotyping algorithms.
- To develop phenotyping algorithms in an unbiased manner, comparable in accuracy to expert-curated features.
- To improve the efficiency of high-throughput phenotyping for research.
Main Methods:
- Collected comprehensive medical concepts from public knowledge sources automatically.
- Utilized natural language processing (NLP) to identify concept occurrence patterns in EHR narratives.
- Selected informative features for phenotype classification and trained a penalized logistic regression model.
Main Results:
- Applied the method to identify rheumatoid arthritis (RA) and coronary artery disease (CAD) cases in EHR data.
- Achieved high classification accuracy with automated features: AUC of 0.951 for RA and 0.929 for CAD.
- Performance was comparable or slightly superior to models trained with expert-curated features.
Conclusions:
- Automated NLP-based feature selection yields phenotyping algorithms with high accuracy and efficiency.
- The developed method offers a significant advancement for high-throughput phenotyping.
- The majority of selected features were interpretable, facilitating clinical understanding.
Related Concept Videos
Genetic Screens
5.9K
Genetic screens are tools used to identify genes and mutations responsible for phenotypes of interest. Genetic screens help identify individuals or a group of people at risk of developing genetic diseases and help them with early intervention, targeted therapy, and reproductive options.
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
5.9K
Light Acquisition
9.9K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
9.9K

