Related Experiment Video
Updated: Aug 9, 2025

Reusable Single Cell for Iterative Epigenomic Analyses
Published on: February 11, 2022
Pitfalls and opportunities for applying latent variables in single-cell eQTL analyses
Angli Xue1,2, Seyhan Yazar3, Drew Neavin3
1Garvan-Weizmann Centre for Cellular Genomics, Garvan Institute of Medical Research, Sydney, NSW, 2010, Australia. a.xue@garvan.org.au.
Latent variables like Probabilistic Estimation of Expression Residuals (PEER) and Principal Component Analysis (PCA) improve single-cell expression quantitative trait loci (eQTL) detection. Validating these factors on pseudo-bulk data enhances eGene discovery across cell types.
Area of Science:
- Genomics
- Computational Biology
- Single-cell analysis
Background:
- Latent variables correct confounders and boost statistical power in gene expression studies.
- Probabilistic Estimation of Expression Residuals (PEER) and Principal Component Analysis (PCA) are common for bulk RNA-seq, but their single-cell application is less understood.
- Single-cell RNA-seq data presents unique challenges like sparsity and skewness, potentially impacting latent variable performance.
Purpose of the Study:
- To evaluate the effectiveness of PEER and PCA for single-cell expression quantitative trait loci (eQTL) detection.
- To identify optimal data processing steps for latent variable generation in single-cell eQTL analysis.
- To assess the impact of cell type heterogeneity on latent variable performance and eGene discovery.
Main Methods:
- Analysis of pseudo-bulk matrices derived from single-cell RNA-seq data.
- Application and validation of PEER and PCA for latent variable identification.
- Integration of latent variables into eQTL association models.
- Sensitivity analyses across different cell types and gene subsets.
Main Results:
- PEER and PCA require specific quality control and transformation steps on pseudo-bulk data to yield valid latent variables, avoiding highly correlated factors.
- Incorporating validated latent variables increased expression-associated gene (eGene) detection by 1.7% to 13.3%.
- Using highly variable genes for latent variable generation achieved comparable eGene discovery to using all genes, with a ~6.2-fold speedup in computation.
Conclusions:
- Properly processed latent variables significantly enhance eQTL discovery in single-cell data.
- Cell type-specific adjustments are crucial for optimizing latent variable-based eQTL analysis.
- Leveraging highly variable genes offers an efficient strategy for latent variable computation without compromising discovery power.
More Related Videos
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
09:20Single-Cell Factor Localization on Chromatin using Ultra-Low Input Cleavage Under Targets and Release using Nuclease
Published on: February 1, 2022