Related Experiment Video
Updated: Feb 19, 2026

Identifying Amino Acid Overproducers Using Rare-Codon-Rich Markers
Published on: June 24, 2019
Pichia-CLM: A language model-based codon optimization pipeline for Komagataella phaffii
Harini Narayanan1, J Christopher Love1,2
1Koch Institute for Integrative Cancer Research, Massachusetts Institute of Technology, Cambridge, MA 02139.
Abstract:
The preference in synonymous codon usage-the so-called codon usage bias (CUB)-is governed by several factors such as the host organism, context and function of the gene, and the position of the codon within the gene itself. We demonstrated that this mapping can be learned from the host's genome using language models and subsequently applied for codon optimization of heterologous proteins expressed by the host. This pipeline called Pichia-Codon language model (Pichia-CLM) was applied to the industrial host organism, Komagataella phaffii. With this approach, production of heterologous proteins was enhanced up to threefold compared to their native sequences. Furthermore, Pichia-CLM consistently yielded constructs with enhanced productivity for proteins of varied complexity, compared to commercially available tools. Finally, we showed that Pichia-CLM generates sequences resembling the properties of codon usage found in the host's intrinsic host cell proteins and learned features such as avoiding negative cis-regulatory and repeat elements based on patterns in the genome data. These results show the potential of language models to unbiasedly learn patterns and design robust sequences for improved protein production.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Leaky Scanning
Translation in Prokaryotes
Point and Frameshift Mutations
From DNA to Protein

