Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multi-species Conserved Sequences02:51

Multi-species Conserved Sequences

4.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale  studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.9K
Peptide Identification Using Tandem Mass Spectrometry01:33

Peptide Identification Using Tandem Mass Spectrometry

8.7K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
8.7K
RNA-seq03:21

RNA-seq

12.2K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.2K
Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

7.1K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.1K
Sanger Sequencing01:57

Sanger Sequencing

776.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
776.0K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

18.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
18.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Evaluating search engines and large language models for answering health questions.

NPJ digital medicine·2025
Same author

Efficient phylogenetic tree inference for massive taxonomic datasets: harnessing the power of a server to analyze 1 million taxa.

GigaScience·2024
Same author

BigSeqKit: a parallel Big Data toolkit to process FASTA and FASTQ files at scale.

GigaScience·2023
Same author

A machine learning approach to model the impact of line edge roughness on gate-all-around nanowire FETs while reducing the carbon footprint.

PloS one·2023
Same author

Big Data in metagenomics: Apache Spark vs MPI.

PloS one·2020
Same author

A Big Data Platform for Real Time Analysis of Signs of Depression in Social Media.

International journal of environmental research and public health·2020

Related Experiment Video

Updated: Mar 1, 2026

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
09:14

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens

Published on: June 28, 2018

7.6K

PASTASpark: multiple sequence alignment meets Big Data.

José M Abuín1, Tomás F Pena1, Juan C Pichel1

  • 1CiTIUS, Universidade de Santiago de Compostela, 15782 Santiago de Compostela, Spain.

Bioinformatics (Oxford, England)
|June 6, 2017
PubMed
Summary

PASTASpark accelerates multiple sequence alignment by leveraging Apache Spark, achieving up to 10x speedups. This enables processing of ultra-large biological datasets within practical time limits.

More Related Videos

A Practical Guide to Phylogenetics for Nonexperts
12:00

A Practical Guide to Phylogenetics for Nonexperts

Published on: February 5, 2014

36.2K
An Integrated Approach for Microprotein Identification and Sequence Analysis
09:37

An Integrated Approach for Microprotein Identification and Sequence Analysis

Published on: July 12, 2022

4.0K

Related Experiment Videos

Last Updated: Mar 1, 2026

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
09:14

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens

Published on: June 28, 2018

7.6K
A Practical Guide to Phylogenetics for Nonexperts
12:00

A Practical Guide to Phylogenetics for Nonexperts

Published on: February 5, 2014

36.2K
An Integrated Approach for Microprotein Identification and Sequence Analysis
09:37

An Integrated Approach for Microprotein Identification and Sequence Analysis

Published on: July 12, 2022

4.0K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Big Data Analytics

Background:

  • Multiple sequence alignment is fundamental in bioinformatics.
  • PASTA is a state-of-the-art tool for multiple sequence alignment.
  • PASTA's performance is limited by shared memory systems.

Purpose of the Study:

  • Introduce PASTASpark, a novel tool for enhanced multiple sequence alignment.
  • Boost the performance of PASTA's computationally intensive alignment phase.
  • Enable processing of large-scale sequence datasets.

Main Methods:

  • Utilized the Big Data engine Apache Spark.
  • Integrated Apache Spark with the PASTA alignment tool.
  • Developed PASTASpark as an Open Source tool.

Main Results:

  • Achieved speedups of up to 10x compared to single-threaded PASTA.
  • Demonstrated the capability to process datasets with 200,000 sequences within 24 hours.
  • Successfully addressed the time-consuming nature of the alignment phase.

Conclusions:

  • PASTASpark significantly enhances the scalability and efficiency of multiple sequence alignment.
  • The tool facilitates the analysis of ultra-large biological sequence datasets.
  • PASTASpark offers a practical solution for computationally demanding alignment tasks.