Related Experiment Video
Updated: Jul 19, 2025

08:50
Predictive Immune Modeling of Solid Tumors
Published on: February 25, 2020
7.0K
Building an optimal predictive model for imputing tissue-specific gene expression by combining genotype and
Sunwoo Jung1, Cue Hyunkyu Lee2, Jae Hoon Sul3
1Interdisciplinary Program in Bioengineering, Seoul National University, Seoul, Republic of Korea.
HGG Advances
|August 14, 2023
Summary
Combining genotype and whole-blood expression data improves tissue-specific gene expression imputation. A merged model, upweighting whole-blood transcriptome data, showed superior performance over standalone models for complex trait analysis.
Area of Science:
- Genomics
- Systems Biology
- Bioinformatics
Background:
- Accurate imputation of tissue-specific gene expression aids in understanding complex human traits.
- Current imputation methods utilize either genotype or whole-blood expression data, both obtainable from blood samples.
Purpose of the Study:
- To develop an optimal predictive model for tissue-specific gene expression imputation by integrating genotype and whole-blood expression data.
- To evaluate and compare the performance of standalone and combined imputation models across 47 human tissues.
Main Methods:
- Evaluated standalone genotype (GEN) and whole-blood expression (WBE) imputation models.
- Developed and tested combined models, including merged datasets (MERGED) and inverse variance-weighted (IVW) integration.
- Optimized a MERGED model by upweighting whole-blood transcriptome data using a fixed ratio of regularization penalty factors.
Main Results:
- The WBE model significantly outperformed the GEN model across most tissues.
- A specific MERGED model demonstrated noticeable improvement over standalone models.
- The optimal MERGED model leveraged a strategic weighting of whole-blood expression data.
Conclusions:
- Combining genotype and whole-blood expression data can enhance tissue-specific gene expression imputation.
- The effectiveness of combined imputation strategies is highly dependent on the chosen integration method and data weighting.

