Related Experiment Video
Updated: Jul 10, 2025

10:58
Protein WISDOM: A Workbench for In silico De novo Design of BioMolecules
Published on: July 25, 2013
17.1K
GENERALIST: A latent space based generative model for protein sequence families
Hoda Akl1, Brooke Emison2, Xiaochuan Zhao1
1Department of Physics, University of Florida, Gainesville, Florida, United States of America.
Plos Computational Biology
|November 27, 2023
Summary
We developed GENERALIST, a novel generative model for protein sequences. It accurately models protein families, even with limited data, and generates realistic sequence ensembles.
Area of Science:
- Computational biology
- Protein engineering
- Bioinformatics
Background:
- Generative models are crucial for protein science but struggle with large proteins and low-coverage families.
- Existing methods face challenges in inference, accuracy, and overfitting.
Purpose of the Study:
- To introduce GENERALIST, a new generative model for protein sequences.
- To address limitations of current generative approaches in protein modeling.
Main Methods:
- GENERALIST employs a nonlinear tensor factorization approach.
- The model is designed for ease of learning, tunability, and accuracy.
Main Results:
- GENERALIST accurately captures high-order amino acid covariation statistics.
- It predicts stable protein structures and generates sequence ensembles closely matching natural ones.
- The model creates an informative latent space for protein sequences.
Conclusions:
- GENERALIST offers an accurate and efficient method for modeling protein sequence variability.
- It overcomes key limitations of existing generative models.
- This tool will advance the study of protein sequence diversity and function.
More Related Videos
Related Concept Videos
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K
Gene Families
8.8K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.8K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains
3.1K
3.1K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K

