Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Protein Families02:47

Protein Families

15.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
15.6K
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

11.1K
Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.1K
Conservation of Protein Domains02:26

Conservation of Protein Domains

3.2K
3.2K
Protein Networks02:26

Protein Networks

4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Gene Families01:57

Gene Families

8.9K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.9K
Protein Organization01:24

Protein Organization

6.8K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
6.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Accurate <i>ab initio</i> gene prediction in eukaryotes with Tiberius in multiple clades.

bioRxiv : the preprint server for biology·2026
Same author

Genomic map of the functionally extinct northern white rhinoceros (<i>Ceratotherium simum cottoni</i>).

Proceedings of the National Academy of Sciences of the United States of America·2025
Same author

Tiberius: end-to-end deep learning with an HMM for gene prediction.

Bioinformatics (Oxford, England)·2024
Same author

learnMSA2: deep protein multiple alignments with large language and hidden Markov models.

Bioinformatics (Oxford, England)·2024
Same author

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA.

Genome research·2024
Same author

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS and TSEBRA.

bioRxiv : the preprint server for biology·2023

Related Experiment Video

Updated: Aug 20, 2025

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

7.1K

learnMSA: learning and aligning large protein families.

Felix Becker1, Mario Stanke1

  • 1Institute of Mathematics and Computer Science, University of Greifswald, Walther-Rathenau-Straße 47, 17489 Greifswald, Germany.

Gigascience
|November 18, 2022
PubMed
Summary

learnMSA, a new method for protein sequence alignment, accurately aligns millions of sequences faster than current tools. This approach improves accuracy with more data, unlike many existing algorithms.

Keywords:
machine learningmultiple sequence alignmentprofile hidden Markov model

More Related Videos

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
07:08

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues

Published on: July 14, 2015

7.4K
An Integrated Approach for Microprotein Identification and Sequence Analysis
09:37

An Integrated Approach for Microprotein Identification and Sequence Analysis

Published on: July 12, 2022

3.5K

Related Experiment Videos

Last Updated: Aug 20, 2025

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

7.1K
Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
07:08

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues

Published on: July 14, 2015

7.4K
An Integrated Approach for Microprotein Identification and Sequence Analysis
09:37

An Integrated Approach for Microprotein Identification and Sequence Analysis

Published on: July 12, 2022

3.5K

Area of Science:

  • Computational Biology
  • Bioinformatics
  • Machine Learning

Background:

  • Protein sequence alignment is crucial for biological data analysis.
  • Current algorithms struggle with accuracy as the number of sequences increases.
  • Inaccurate alignments hinder downstream biological tasks.

Purpose of the Study:

  • To develop a novel, accurate, and scalable method for multiple sequence alignment.
  • To address the limitations of existing algorithms in handling large biological datasets.
  • To improve the accuracy and speed of protein sequence alignment.

Main Methods:

  • Implemented a statistical learning approach using profile hidden Markov models (pHMMs) and batch gradient descent.
  • Utilized a custom recurrent neural network architecture for pHMMs, trained with maximum a posteriori objective.
  • Employed automatic differentiation and uniform batch sampling for efficient training on large datasets without progressive heuristics.

Main Results:

  • learnMSA demonstrated superior accuracy and speed on ultra-large protein families (up to 3.5 million sequences).
  • The method achieved state-of-the-art performance on established benchmarks (HomFam, BaliFam).
  • Experiments were conducted on a standard workstation with GPU acceleration.

Conclusions:

  • learnMSA overcomes the accuracy limitations of heuristic aligners when dealing with large numbers of sequences.
  • The framework offers a scalable and accurate solution for large-scale multiple sequence alignment.
  • Future improvements and applications of learnMSA are anticipated.