Related Experiment Video
Updated: Jul 8, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Bayesian Double Feature Allocation for Phenotyping with Electronic Health Records.
Yang Ni1,2, Peter Müller3, Yuan Ji4
1Department of Statistics, Texas A&M University.
We developed a new statistical method using electronic health records to find hidden diseases. This approach identifies 10 distinct latent diseases, aiding in disease prevention and health monitoring.
Area of Science:
- Computational biology
- Statistical genetics
- Health informatics
Background:
- Electronic health records (EHR) offer valuable data for understanding complex human phenotypes.
- Statistical modeling of EHR data can reveal underlying disease patterns and patient subgroups.
Purpose of the Study:
- To propose a novel categorical matrix factorization method for inferring latent diseases from EHR data.
- To enhance the identifiability and interpretability of latent diseases using Bayesian approaches and prior knowledge of known diseases.
Main Methods:
- A double feature allocation model was developed to simultaneously allocate features to rows and columns of a categorical matrix.
- Bayesian inference was employed, incorporating prior information on known diseases like hypertension and diabetes.
- The method was validated through simulation studies and compared with sparse latent factor models.
Main Results:
- Application to a Chinese EHR dataset identified 10 latent diseases.
- These latent diseases were associated with specific health traits including lipid disorders, thrombocytopenia, polycythemia, anemia, infections, allergy, and malnutrition.
- The identified latent diseases showed agreement with findings reported in medical literature.
Conclusions:
- The proposed method effectively infers latent diseases from EHR data, offering insights into complex health conditions.
- This approach can assist healthcare officials in monitoring patient health, identifying risk factors, and developing preventive strategies.
- An R package ('dfa') and a web application are available for implementing the method and exploring the findings.
More Related Videos
05:53Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
09:38Generalized Psychophysiological Interaction PPI Analysis of Memory Related Connectivity in Individuals at Genetic Risk for Alzheimer's Disease
Published on: November 14, 2017
Related Concept Videos
Multiple Allele Traits
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Probability Laws
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.