Extending Protein Language Models to a Viral Genomic Scale Using Biologically Induced Sparse Attention

Thibaut Dejean1, Barbra D Ferrell2, Zachary D Schreiber2

  • 1Department of Information and Computer Sciences University of Hawaii Honolulu, HI.

Gigascience
|July 21, 2026
PubMed
Summary

This study introduces a novel long-context protein language model trained on entire viral genomes. This approach captures inter-protein relationships, improving predictions for masked amino acids and downstream tasks.

Related Concept Videos

Leaky Scanning02:28

Leaky Scanning

During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA.  Marilyn Kozak discovered that the sequence RCCAUGG (where R stands for...
Size and Structure of Viral Genomes01:26

Size and Structure of Viral Genomes

Viral genomes exhibit remarkable diversity in size, structure, and composition, influencing their replication strategies and interactions with host cells. These genomes consist of either DNA or RNA and may be linear or circular. Additionally, they can be single-stranded or double-stranded, with each configuration affecting how the virus propagates within a host. RNA viruses, for instance, generally have smaller genomes than DNA viruses, a factor that contributes to their high mutation rates and...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Viruses with RNA Genomes01:29

Viruses with RNA Genomes

RNA viruses are categorized into positive-strand, negative-strand, or double-stranded groups based on their genomic structure and replication mechanisms. This classification dictates how they exploit host cellular machinery for protein synthesis and replication. Some RNA viruses also utilize reverse transcription as part of their life cycle, further diversifying their replication strategies.Positive-Strand RNA VirusesPositive-strand RNA viruses have genomes that function directly as messenger...