Related Experiment Video
Updated: Jun 9, 2026

07:09
A Bioinformatics Pipeline for Investigating Molecular Evolution and Gene Expression using RNA-seq
Published on: May 28, 2021
HyDRA: A pipeline for integrating long- and short-read RNAseq data for custom transcriptome assembly
Isabela Almeida1,2, Xue Lu1, Stacey L Edwards1,2,3
1Cancer Program, QIMR Berghofer, Brisbane, QLD 4029, Australia.
Iscience
|June 8, 2026
Summary
This study introduces HyDRA, a novel pipeline for transcriptome assembly. HyDRA combines short and long sequencing reads to accurately reconstruct full-length transcripts and discover novel noncoding RNAs.
Area of Science:
- Genomics
- Bioinformatics
- Transcriptomics
Background:
- Short-read RNA sequencing (RNAseq) is limited in reconstructing full-length transcripts and capturing transcript diversity.
- Long-read RNAseq offers structural resolution but suffers from high error rates.
- Noncoding RNA transcripts are often underrepresented in current genomic references.
Purpose of the Study:
- To develop a hybrid bioinformatics pipeline for *de novo* transcriptome assembly.
- To integrate short-read accuracy with long-read structural resolution.
- To improve the completeness and accuracy of transcriptome reconstruction, particularly for noncoding RNAs.
Main Methods:
- Presentation of the hybrid *de novo* RNA assembly (HyDRA) pipeline.
- Integration of short-read and long-read sequencing data.
- Benchmarking of HyDRA against existing transcriptome assembly methods.
Main Results:
- HyDRA outperforms existing methods by up to 40% in transcriptome assembly.
- Identification of over 50,000 high-confidence long noncoding RNAs in the human ovarian metatranscriptome.
- Discovery of numerous previously undetected noncoding RNAs.
Conclusions:
- HyDRA provides a more complete *de novo* transcriptome assembly by leveraging both short and long reads.
- The pipeline significantly enhances the detection of long noncoding RNAs.
- HyDRA is crucial for advancing the understanding of transcriptomic complexity and the noncoding genome, especially given the abundance of short-read data.
