Related Experiment Video
Updated: Jul 24, 2025

12:08
Hybrid De Novo Genome Assembly for the Generation of Complete Genomes of Urinary Bacteria using Short- and Long-read Sequencing Technologies
Published on: August 20, 2021
5.1K
From contigs towards chromosomes: automatic improvement of long read assemblies (ILRA)
José Luis Ruiz1, Susanne Reimering2, Juan David Escobar-Prieto3
1Instituto de Parasitología y Biomedicina López-Neyra (IPBLN), Consejo Superior de Investigaciones Científicas, 18016, Granada, Spain.
Briefings in Bioinformatics
|July 5, 2023
Summary
The ILRA pipeline improves long-read genome assemblies by correcting errors and enhancing contiguity. This tool aids researchers in generating high-quality genome sequences more efficiently.
Area of Science:
- Genomics
- Bioinformatics
Background:
- Long-read sequencing technologies offer advancements for genome assembly but present challenges with repeats and homopolymer errors.
- Existing methods often result in numerous contigs and inaccuracies, hindering complete genome reconstruction.
Purpose of the Study:
- To introduce and evaluate the ILRA (Improved Long Read Assembly) pipeline for correcting errors in long-read genome assemblies.
- To enhance the contiguity and accuracy of genome sequences generated using long-read technologies.
Main Methods:
- The ILRA pipeline first reorders, renames, merges, circularizes, or filters contigs from long-read data.
- Illumina short reads are then utilized to correct homopolymer indel errors within the assemblies.
- The pipeline's performance was benchmarked on various species, including Homo sapiens, Trypanosoma brucei, Leptosphaeria spp., and Plasmodium falciparum.
Main Results:
- ILRA successfully improved the quality of long-read genome assemblies up to 1 Gbp.
- Correction of homopolymer tracts reduced the misannotation of genes as pseudogenes.
- The study identified that an iterative application of the pipeline may be necessary for further error correction.
Conclusions:
- The ILRA pipeline is an effective tool for enhancing the quality of long-read genome assemblies.
- This approach addresses key limitations of long-read sequencing, improving genome sequence accuracy and contiguity.
- The pipeline is publicly available, facilitating its adoption in genomic research.
Related Concept Videos
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K
RNA-seq
10.1K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.1K

