Related Experiment Video
Updated: Aug 12, 2025

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
3.5K
Large language models generate functional protein sequences across diverse families
Ali Madani1,2, Ben Krause3, Eric R Greene4
1Salesforce Research, Palo Alto, CA, USA. ali@madani.ai.
Nature Biotechnology
|January 26, 2023
Summary
ProGen, a novel language model, generates functional protein sequences across diverse families. This artificial intelligence approach aids protein design and engineering by mimicking natural protein capabilities.
Area of Science:
- Biotechnology
- Computational Biology
- Protein Engineering
Background:
- Deep-learning language models show potential in protein design and engineering.
- Generating functional protein sequences across large families remains a challenge.
Purpose of the Study:
- Introduce ProGen, a language model for generating protein sequences with predictable functions.
- Demonstrate ProGen's adaptability and performance in protein family generation.
Main Methods:
- Trained ProGen on 280 million protein sequences from over 19,000 families.
- Augmented the model with control tags for specifying protein properties.
- Fine-tuned ProGen on curated sequences and tags for improved controllable generation.
Main Results:
- ProGen successfully generated functional protein sequences across diverse families, including lysozymes, chorismate mutase, and malate dehydrogenase.
- Artificial lysozymes engineered by ProGen exhibited catalytic efficiencies comparable to natural lysozymes.
- Achieved significant sequence identity reduction (as low as 31.4%) while maintaining protein function.
Conclusions:
- ProGen offers a powerful tool for de novo protein design and engineering.
- The model's ability to generate functional proteins with high control advances biotechnological applications.
- ProGen demonstrates broad applicability across various protein families.
Related Concept Videos
Protein Families
15.5K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.5K
Conservation of Protein Domains Over Different Proteins
11.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.0K
Evolutionary Relationships through Genome Comparisons
6.1K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.1K
Conserved Binding Sites
4.3K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.3K
Multi-species Conserved Sequences
4.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.0K
Conservation of Protein Domains
3.2K
3.2K

