Related Experiment Video
Updated: Oct 7, 2025

08:35
Identification of Alternative Splicing and Polyadenylation in RNA-seq Data
Published on: June 24, 2021
5.9K
Optimized splitting of mixed-species RNA sequencing data
Xuan Song1, Hai Yun Gao2, Karl Herrup1
1Department of Neurology, Alzheimer's Disease Research Center, University of Pittsburgh, Pittsburgh, PA 15213, USA.
Journal of Bioinformatics and Computational Biology
|January 7, 2022
Summary
Analyzing mixed-species RNA sequencing data is crucial for biological research. This study compares alignment-dependent and alignment-independent methods, finding that optimized alignment offers higher accuracy for transcript quantification in xenograft and co-culture models.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Gene expression studies using mixed-species models (e.g., human and mouse xenografts/co-cultures) are vital for understanding cellular dynamics in development and disease.
- Challenges in analyzing mixed-species RNA sequencing data arise from similar mRNA sequences between species, impacting accurate transcript quantification.
Purpose of the Study:
- To identify optimal strategies for analyzing mixed-species RNA sequencing data.
- To evaluate the performance of alignment-dependent and alignment-independent methods for species-specific read classification.
Main Methods:
- Evaluation of alignment-dependent methods, including alignment to a pooled reference index followed by re-alignment to individual genomes.
- Assessment of alignment-independent methods, such as convolutional neural networks, for classifying RNA sequencing reads based on conserved sequence patterns.
Main Results:
- Alignment to a pooled reference index, with optimal alignment for species classification and subsequent re-alignment to individual genomes, achieved high accuracy across various species ratios.
- Alignment-independent methods, like convolutional neural networks, demonstrated over 85% accuracy in classifying RNA sequencing reads.
- Both approaches performed well with different human and mouse read ratios, but mixed-genome alignment with optimized read separation showed lower error rates.
Conclusions:
- Optimized alignment-based strategies provide a more accurate approach for transcript quantification in mixed-species RNA sequencing data compared to non-alignment methods.
- Effective species-specific read partitioning is achievable with both alignment-dependent and alignment-independent methods, but traditional alignment offers superior precision.
Related Concept Videos
RNA-seq
10.5K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.5K
Multi-species Conserved Sequences
4.3K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.3K

