Related Experiment Video
Updated: Jul 4, 2025

Curation of Computational Chemical Libraries Demonstrated with Alpha-Amino Acids
Published on: April 13, 2022
Protein language models meet reduced amino acid alphabets
Ioan Ieremie1, Rob M Ewing2, Mahesan Niranjan1
1Vision, Learning & Control Group, University of Southampton, Southampton SO17 1BJ, United Kingdom.
Protein language models (PLMs) trained on reduced amino acid alphabets show that full alphabets capture more detail. However, reduced alphabets can improve protein structure prediction accuracy in some cases.
Area of Science:
- Computational Biology
- Bioinformatics
- Protein Science
Background:
- Protein language models (PLMs) leverage natural language processing techniques for unsupervised representation learning.
- PLMs have significantly improved performance in various downstream protein-related tasks.
- Previous research explored reduced amino acid alphabets based on physicochemical properties, but their use in PLMs and folding models remains underexplored.
Purpose of the Study:
- To evaluate the effectiveness of PLMs trained on reduced amino acid alphabets.
- To understand how information loss from alphabet reduction affects learned protein representations and downstream task performance.
- To assess the impact of reduced alphabets on protein structure prediction using ESMFold.
Main Methods:
- Training protein language models on datasets with reduced amino acid alphabets.
- Comparing the representational capacity of PLMs trained on full versus reduced alphabets.
- Utilizing ESMFold to predict structures of proteins translated into reduced alphabets.
- Analyzing the impact of reduced alphabets on structural prediction accuracy (LDDT-Cα).
Main Results:
- PLMs trained on the full amino acid alphabet and extensive sequence data capture finer details compared to reduced alphabet methods.
- Protein structure prediction using ESMFold with reduced alphabets showed improved results for 10 out of 50 CASP14 targets.
- Structural prediction accuracy improvements reached up to 19% in LDDT-Cα differences for specific proteins.
Conclusions:
- While full alphabets in PLMs retain more detailed evolutionary information, reduced alphabets offer a potential avenue for enhancing protein structure prediction accuracy in specific instances.
- The findings highlight a trade-off between information richness and potential improvements in downstream applications like structure prediction when using reduced amino acid alphabets.
More Related Videos
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
11:47Residue-specific Incorporation of Noncanonical Amino Acids into Model Proteins Using an Escherichia coli Cell-free Transcription-translation System
Published on: August 1, 2016
Related Concept Videos
From DNA to Protein
Amino acids
tRNA Activation
What are Proteins?
The Central Dogma
Protein Organization