Related Experiment Video
Updated: Jan 11, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
A Scalable Framework to Integrate Social Determinants of Health into Disease Risk Models using Biobank Survey Data
Jane Brown1,2, Abhijith Biji1,2, Kathleen Ferar1,2
1Institute for Genomic Health, Icahn School of Medicine at Mount Sinai, New York, NY.
Predicting complex diseases is challenging. This study uses Multiple Correspondence Analysis (MCA) on biobank data to create non-genetic risk profiles, significantly improving disease prediction beyond genetics alone.
Area of Science:
- Genomics and Epidemiology
- Computational Biology
- Public Health
Background:
- Complex diseases pose a significant global health challenge, with limited ability to predict individual risk.
- Disease risk is influenced by a combination of genetic and non-genetic factors (environmental, behavioral, social), which are seldom analyzed together.
- Large-scale biobanks offer opportunities to integrate diverse data for improved disease risk prediction.
Purpose of the Study:
- To develop and validate a scalable framework using Multiple Correspondence Analysis (MCA) to model non-genetic determinants of complex diseases.
- To assess the contribution of non-genetic factors, summarized by MCA embeddings, to disease risk prediction.
- To investigate the interplay between genetic and non-genetic risk factors in complex diseases.
Main Methods:
- Applied Multiple Correspondence Analysis (MCA) to over 100 environmental, behavioral, and social variables from the All of Us biobank (N=171,614).
- Generated low-dimensional embeddings to quantify non-genetic risk for six common chronic conditions.
- Integrated MCA embeddings with demographic data and polygenic scores (PGS) to evaluate prediction performance (ROC-AUC).
Main Results:
- MCA embeddings identified known and novel risk factors for chronic diseases.
- Non-genetic risk embeddings consistently improved prediction accuracy beyond demographics and PGS (ROC-AUC increase: 0.03-0.05).
- For five of six diseases, MCA embeddings outperformed PGS in prediction improvement; genetic and non-genetic risks largely acted additively.
Conclusions:
- Introduced a scalable, interpretable MCA framework for summarizing complex survey-based non-genetic factors.
- Demonstrated that non-genetic factors substantially enhance complex disease risk prediction, acting largely additively with genetic risk.
- Clarified gene-environment interplay, supporting more equitable and robust risk modeling for diverse populations.
More Related Videos
07:44Author Spotlight: Implementation of BIVA for Analyzing Disease Risk Factors in Patients with Low Body Cell Mass
Published on: July 14, 2023
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Overview of Biostatistics in Health Sciences
Statistical Software for Data Analysis and Clinical Trials
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Biostatistics: Overview
Discrete variables are...
Bias in Epidemiological Studies