Intrinsic dataset features drive mutational effect prediction by protein language models.

Luiz C Vieira1, Sophia Lin1, Claus O Wilke1

  • 1Department of Integrative Biology, The University of Texas at Austin, Austin, TX, United States of America.

Summary

Protein language models (pLMs) show variable performance in predicting protein fitness landscapes. Dataset composition, particularly site variability, is more critical than model architecture for accurate mutational effect prediction.

Related Concept Videos

Mutations01:39

Mutations

Overview
98.0K
Mutations01:35

Mutations

Mutations are changes in the sequence of DNA. These changes can occur spontaneously or they can be induced by exposure to environmental factors. Mutations can be characterized in a number of different ways: whether and how they alter the amino acid sequence of the protein, whether they occur over a small or large area of DNA, and whether they occur in somatic cells or germline cells.
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
45.7K
Covalently Linked Protein Regulators02:04

Covalently Linked Protein Regulators

Proteins can undergo many types of post-translational modifications, often in response to changes in their environment. These modifications play an important role in the function and stability of these proteins. Covalently linked molecules include functional groups, such as methyl, acetyl, and phosphate groups, and also small proteins, such as ubiquitin. There are around 200 different types of covalent regulators that have been identified.
These groups modify specific amino acids in a protein....
9.9K
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.6K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.8K
Translation01:31

Translation

Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Proteins are...
23.2K