Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

10.7K
Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.7K
From DNA to Protein03:06

From DNA to Protein

17.9K
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
17.9K
Conservation of Protein Domains02:26

Conservation of Protein Domains

3.1K
3.1K
Leaky Scanning02:28

Leaky Scanning

5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA.  Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Protein Organization01:24

Protein Organization

6.2K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
6.2K
Protein Families02:47

Protein Families

15.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
15.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

On the state of protein function prediction: a report on the fourth CAFA challenge.

bioRxiv : the preprint server for biology·2026
Same author

Advances in Protein Function Prediction from the Fifth CAFA Challenge.

bioRxiv : the preprint server for biology·2026
Same author

Whole-genome prediction of bacterial pathogenic capacity on novel bacteria using protein language models with PathogenFinder2.

Bioinformatics (Oxford, England)·2026
Same author

Biocentral: Embedding-based Protein Predictions.

Journal of molecular biology·2026
Same author

Toxin data quality: a critical examination of bacterial exotoxins and animal toxins.

BMC research notes·2025
Same author

FlatProt: 2D visualization eases protein structure comparison.

BMC bioinformatics·2025

Related Experiment Video

Updated: May 28, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

478

Are protein language models the new universal key?

Konstantin Weissenow1, Burkhard Rost2

  • 1TUM (Technical University of Munich), School of Computation, Information and Technology (CIT), Faculty of Informatics, Chair of Bioinformatics & Computational Biology - i12, Boltzmannstr. 3, 85748 Garching/Munich, Germany; TUM Graduate School, Center of Doctoral Studies in Informatics and its Applications (CeDoSIA), Boltzmannstr. 11, 85748 Garching, Germany.

Current Opinion in Structural Biology
|February 8, 2025
PubMed
Summary

Protein language models (pLMs) offer a powerful new approach to protein prediction, surpassing traditional methods in accuracy and efficiency. These models, using embeddings, provide protein-specific insights with fewer computational resources.

Keywords:
genome sequence analysismultiple alignmentspredicting globularityprotein domainsprotein structure predictionsolvent accessibilitytransmembrane helices

More Related Videos

Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning
05:33

Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning

Published on: January 29, 2020

5.9K
A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

68.4K

Related Experiment Videos

Last Updated: May 28, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

478
Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning
05:33

Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning

Published on: January 29, 2020

5.9K
A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

68.4K

Area of Science:

  • Computational Biology
  • Bioinformatics
  • Machine Learning in Biology

Background:

  • Traditional protein prediction relied heavily on evolutionary information from Multiple Sequence Alignments (MSAs).
  • Protein Language Models (pLMs) learn the inherent grammar of protein sequences.
  • pLM embeddings implicitly encode this learned grammatical information.

Purpose of the Study:

  • To evaluate the efficacy of pLM embeddings as exclusive inputs for protein prediction tasks.
  • To compare the performance and resource efficiency of pLM-based methods against traditional MSA-based approaches.
  • To advocate for optimizing existing pLM foundation models over retraining new ones.

Main Methods:

  • Utilizing pLM-generated embeddings as the sole input for downstream supervised learning models.
  • Comparing prediction accuracy and computational resource consumption with established MSA-based methods.
  • Assessing the efficiency of pLM embeddings in condensing complex protein sequence information.

Main Results:

  • MSA-free pLM-based predictions demonstrate significantly improved accuracy across many applications.
  • pLM embeddings efficiently condense sequence grammar, enabling downstream models with fewer parameters.
  • pLM-based solutions offer protein-specific predictions and require substantially fewer computational resources post-pretraining.

Conclusions:

  • pLMs are rapidly emerging as a universal and highly effective tool for protein prediction.
  • pLM-based methods present a more accurate and resource-efficient alternative to traditional MSA-based techniques.
  • The study encourages community focus on optimizing pLM foundation models for broader adoption and sustainability.