Related Experiment Video
Updated: Jun 29, 2026

12:54
Real-time Analysis of Transcription Factor Binding, Transcription, Translation, and Turnover to Display Global Events During Cellular Activation
Published on: March 7, 2018
Compound models and Pearson residuals for single-cell RNA-seq data without UMIs
Jan Lause1, Christoph Ziegenhain2, Leonard Hartmanis3
1Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany.
Genome Biology
|June 27, 2026
Summary
This study introduces a novel compound distribution model for normalizing single-cell RNA sequencing (scRNA-seq) data, improving analysis of non-UMI datasets by capturing overdispersion and zero-inflation.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Existing methods for single-cell RNA sequencing (scRNA-seq) data normalization often rely on Pearson residuals from Poisson or negative binomial models, primarily for Unique Molecular Identifier (UMI)-based data.
- Normalization of non-UMI scRNA-seq data presents challenges due to amplification biases and complex distributional patterns.
Purpose of the Study:
- To extend Pearson residual-based normalization methods to non-UMI scRNA-seq data.
- To develop a statistical model that accurately captures the technical noise inherent in non-UMI sequencing protocols.
- To improve gene selection and data representation (embeddings) for Smart-seq2 and similar non-UMI datasets.
Main Methods:
- Modeled the RNA amplification step using a compound distribution, combining a negative binomial distribution for captured molecules with an amplification distribution.
- Derived compound Pearson residuals from this novel model.
- Characterized amplification distributions across various sequencing protocols using a broken power law model.
Main Results:
- The compound Pearson residuals effectively enabled meaningful gene selection and generated informative embeddings for Smart-seq2 datasets.
- Demonstrated that amplification distributions in several sequencing protocols follow a broken power law.
- The developed compound model successfully accounts for previously unaddressed overdispersion and zero-inflation in non-UMI scRNA-seq data.
Conclusions:
- The proposed compound distribution model provides a robust framework for normalizing non-UMI scRNA-seq data.
- This approach enhances the accuracy and interpretability of analyses for datasets generated by protocols like Smart-seq2.
- The findings offer a more comprehensive understanding of technical noise in single-cell genomics data.
Related Concept Videos
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Ribosome Profiling
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...

