Related Experiment Video
Updated: Mar 2, 2026

09:30
Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms
Published on: September 13, 2018
10.0K
Gene length and detection bias in single cell RNA sequencing protocols
Belinda Phipson1, Luke Zappia1,2, Alicia Oshlack1,2
1Murdoch Childrens Research Institute, Parkville, Victoria, 3052, Australia.
F1000Research
|May 23, 2017
Summary
Unique molecular identifiers (UMIs) in single-cell RNA sequencing (scRNA-seq) mitigate gene length bias, unlike full-length protocols. Combining both UMI and full-length data can reveal underlying biological insights in mouse embryonic stem cells.
Area of Science:
- Genomics
- Molecular Biology
- Bioinformatics
Background:
- Single-cell RNA sequencing (scRNA-seq) is crucial for transcriptomic profiling, enabling cell type discovery and tissue development insights.
- Technical challenges in scRNA-seq include amplification biases due to limited starting material.
- Unique molecular identifiers (UMIs) mitigate amplification biases by tagging individual molecules for accurate transcript abundance estimation.
Purpose of the Study:
- To investigate the impact of gene length bias in scRNA-seq across diverse datasets.
- To compare gene length bias between full-length and UMI-based scRNA-seq protocols.
- To assess the potential for combining full-length and UMI data for biological discovery.
Main Methods:
- Analysis of multiple scRNA-seq datasets with varying capture technologies, library preparations, cell types, and species.
- Comparison of gene detection rates and dropout events between full-length and UMI-based protocols.
- Examination of gene length distribution in datasets detected exclusively by either full-length or UMI methods.
Main Results:
- Full-length scRNA-seq protocols exhibit gene length bias, similar to bulk RNA-seq, with shorter genes showing lower counts and higher dropout.
- UMI-based scRNA-seq protocols do not show gene length bias, maintaining uniform dropout rates across gene lengths.
- Genes detected only in UMI datasets were shorter, while those detected only in full-length datasets were longer, using mouse embryonic stem cell data.
Conclusions:
- The choice of scRNA-seq protocol significantly influences gene detection rates and introduces gene length bias in full-length datasets.
- UMI-based protocols effectively eliminate gene length bias, offering more uniform gene detection.
- Combining full-length and UMI data can overcome individual protocol limitations and enhance the understanding of cellular expression patterns.

