Related Experiment Video
Updated: Jul 9, 2026

Protein WISDOM: A Workbench for In silico De novo Design of BioMolecules
Published on: July 25, 2013
Optimizing protein tokenization: reduced amino acid alphabets for efficient and accurate protein language models
1The Shmunis School of Biomedicine and Cancer research, George S. Wise Faculty of Life Sciences, Tel Aviv University, Tel Aviv, 6997801, Israel.
Using reduced amino acid alphabets with Byte Pair Encoding (BPE) tokenization significantly shortens protein language model (pLM) sequences. This approach enhances computational efficiency during training and inference with minimal impact on predictive performance.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning in Biology
Background:
- Protein language models (pLMs) traditionally use a 20-amino-acid alphabet, leading to long sequences and high computational costs.
- Sub-word tokenization (e.g., Byte Pair Encoding/BPE) can shorten sequences but struggles with protein data sparsity.
- Reduced amino acid alphabets group residues by physicochemical properties, offering a potential but understudied solution for tokenization.
Purpose of the Study:
- To investigate the combined efficacy of reduced amino acid alphabets and BPE tokenization in protein language models.
- To assess the impact of this combined approach on sequence length, computational efficiency, and predictive performance.
- To determine if alphabet reduction can improve sub-word tokenization for pLMs.
Main Methods:
- Pre-training RoBERTa-based protein language models (pLMs) de novo.
- Utilizing multiple reduced amino acid alphabets in conjunction with Byte Pair Encoding (BPE) tokenization.
- Evaluating model performance across a diverse range of downstream biological tasks.
Main Results:
- Reduced amino acid alphabets significantly decrease input sequence lengths for pLMs.
- Faster training and inference times were observed when using reduced alphabets.
- Sub-word tokenization with reduced alphabets showed marginal impact on overall predictive performance, with accuracy improvements in specific tasks.
Conclusions:
- Combining reduced amino acid alphabets with BPE tokenization offers a viable strategy to enhance pLM efficiency.
- This approach facilitates more effective sub-word tokenization, reducing computational burden.
- Alphabet reduction presents a promising avenue for optimizing protein language models without compromising predictive accuracy.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
From DNA to Protein
tRNA Activation
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Improving Translational Accuracy
Improving Translational Accuracy

