Related Experiment Video
Updated: Feb 26, 2026

08:23
De novo Identification of Actively Translated Open Reading Frames with Ribosome Profiling Data
Published on: February 18, 2022
4.2K
Advancing codon language modeling with synonymous codon constrained masking.
James Heuschkel1,2, Laura Kingsley1, Noah Pefaur1
1Biotherapeutics Discovery Department, Boehringer Ingelheim Pharmaceutical Inc., Ridgefield, CT 06877, United States.
Nucleic Acids Research
|February 25, 2026
Summary
SynCodonLM, a new codon language model, separates codon and amino acid meanings for better DNA sequence analysis. This model improves understanding of DNA-level biology and protein expression.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Current codon language models often mix codon usage with amino acid semantics, hindering DNA-level biological insights.
- Existing models struggle to capture nucleotide-specific patterns due to conflated semantic information.
Purpose of the Study:
- To introduce SynCodonLM, a novel codon language model designed to disentangle codon-level and protein-level semantics.
- To enable the learning of nucleotide-specific patterns by enforcing biologically grounded constraints.
Main Methods:
- Developed SynCodonLM, a codon language model with a constraint ensuring masked codons are predicted only from synonymous options, guided by the protein sequence.
- Implemented a masking strategy that excludes non-synonymous codons from the prediction space before softmax.
- Modified clustering to group codons by nucleotide properties instead of amino acid identity.
Main Results:
- SynCodonLM successfully disentangles codon and protein semantics, learning nucleotide-specific patterns.
- The model reveals biological structure aligned with DNA-level properties through its clustering approach.
- SynCodonLM outperformed existing models on six out of seven benchmarks sensitive to DNA-level features, including mRNA and protein expression.
Conclusions:
- SynCodonLM represents an advancement in domain-specific representation learning for DNA sequences.
- The model offers new possibilities for sequence design in synthetic biology and deeper exploration of bioprocesses.
- This approach enhances the ability to model DNA sequences by respecting biological constraints.
Related Concept Videos
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
Improving Translational Accuracy
3.7K
3.7K
Leaky Scanning
5.8K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.8K
Per-Unit Sequence Models
469
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
469
Nonsense-mediated mRNA Decay
12.0K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
12.0K
Nonsense-mediated mRNA Decay
3.5K
3.5K

