ACMGA: a reference-free multiple-genome alignment pipeline for plant species
Huafeng Zhou1,2, Xiaoquan Su3, Baoxing Song4,5
1College of Computer Science and Technology, Qingdao University, Qingdao, Shandong, 266071, China.
BMC Genomics
|May 25, 2024
Summary
AnchorWave-Cactus Multiple Genome Alignment (ACMGA) pipeline improves plant genome analysis by identifying more variants than short-read sequencing. This tool enhances the detection of single nucleotide variants and long insertions/deletions, especially in repetitive regions.
Area of Science:
- Genomics
- Bioinformatics
- Plant Science
Background:
- Short-read whole-genome sequencing (WGS) is common for plant genomic variation studies.
- Advancements in long-read sequencing yield high-quality plant genomes.
- Existing multiple genome alignment (MGA) tools may not be optimal for plant genomes.
Purpose of the Study:
- Develop and evaluate a novel MGA pipeline for plant genomes.
- Improve the identification of genomic variants missed by WGS.
- Enhance the alignment of repeat elements and detection of long insertions/deletions (INDELs).
Main Methods:
- Developed the AnchorWave-Cactus Multiple Genome Alignment (ACMGA) pipeline.
- Performed MGA on de novo assembled genomes of Arabidopsis and Maize using ACMGA and Cactus.
- Compared MGA results with previous short-read variant calling data.
Main Results:
- ACMGA improved alignment of repeat elements and identified long INDELs (> 50 bp).
- MGA identified more single nucleotide variants (SNVs) and long INDELs than WGS.
- ACMGA detected significantly more SNVs and INDELs in repetitive regions and overall compared to Cactus.
- ACMGA results showed higher concordance with previously published short-read variants than Cactus results.
- Both MGA pipelines identified numerous multi-allelic variants missed by WGS.
Conclusions:
- Aligning de novo assembled genomes identifies more SNVs and INDELs than short-read mapping.
- ACMGA integrates global alignment, a 2-piece-affine-gap cost strategy, and progressive MGA for effective plant MGA.
- ACMGA offers a practical solution for comprehensive plant genomic variation analysis.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K


