Related Experiment Video
Updated: Apr 19, 2026

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
4.2K
PASTA: Ultra-Large Multiple Sequence Alignment for Nucleotide and Amino-Acid Sequences.
Siavash Mirarab1, Nam Nguyen, Sheng Guo
11 Department of Computer Science, University of Texas at Austin , Austin, Texas.
Summary
We developed PASTA, a novel multiple sequence alignment algorithm, achieving high accuracy and scalability for large biological datasets. PASTA outperforms existing methods in both alignment and phylogenetic tree reconstruction.
Area of Science:
- Computational Biology
- Bioinformatics
- Genomics
Background:
- Accurate multiple sequence alignment (MSA) is crucial for understanding evolutionary relationships and protein function.
- Existing MSA methods face challenges in scalability and accuracy when dealing with large datasets.
Purpose of the Study:
- To introduce PASTA (Partitioning Alignment Search Tree Algorithm), a new MSA algorithm.
- To evaluate PASTA's accuracy, scalability, and performance compared to leading methods.
Main Methods:
- PASTA utilizes a novel guide-tree-based alignment technique.
- The algorithm was tested on biological and simulated datasets, including up to 200,000 sequences.
- Performance was benchmarked against state-of-the-art methods like SATé.
Main Results:
- PASTA demonstrates superior accuracy and scalability in multiple sequence alignment.
- Phylogenetic trees inferred from PASTA alignments show high accuracy, outperforming other methods.
- PASTA is computationally efficient, being faster and requiring less memory than SATé.
Conclusions:
- PASTA represents a significant advancement in multiple sequence alignment algorithms.
- Its high accuracy and scalability make it suitable for analyzing large-scale genomic and proteomic data.
- PASTA facilitates more reliable phylogenetic analyses.
Related Concept Videos
Multi-species Conserved Sequences
5.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
5.0K
Evolutionary Relationships through Genome Comparisons
7.3K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.3K
Modern Molecular Taxonomy
873
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
873
Genome Annotation and Assembly
22.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
22.3K

