Related Experiment Video
Updated: May 13, 2025

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Large-Scale Multi-omic Biosequence Transformers for Modeling Protein-Nucleic Acid Interactions
Sully F Chen1, Robert J Steele2, Glen M Hocky3
1Duke University School of Medicine, Durham, NC 27710, USA.
OmniBioTE, a novel multi-omic model, integrates protein and nucleic acid data, outperforming single-omic approaches. This advancement enhances biomolecule property prediction and reveals emergent structural insights from sequence data alone.
Area of Science:
- Bioinformatics
- Computational Biology
- Molecular Biology
Background:
- Transformer models have advanced biomolecule analysis, but typically focus on single data types (proteins or nucleic acids).
- This single-omic approach limits the models' ability to capture interactions between different biological molecules.
- Existing models struggle to integrate diverse biological sequence data effectively.
Purpose of the Study:
- To introduce OmniBioTE, the largest open-source multi-omic transformer model.
- To demonstrate the capability of a unified model to learn joint representations from mixed protein and nucleic acid sequences.
- To establish a new benchmark for multi-omic sequence analysis and prediction.
Main Methods:
- Trained OmniBioTE on over 250 billion tokens of mixed protein and nucleic acid sequence data.
- Evaluated OmniBioTE's ability to learn representations consistent with the central dogma of molecular biology.
- Assessed performance on predicting binding free energy changes (ΔG) and identifying protein-nucleic acid interaction sites.
Main Results:
- OmniBioTE learned joint representations reflecting the central dogma without explicit biological labels.
- Achieved state-of-the-art results in predicting protein-nucleic acid binding free energy (ΔG).
- Demonstrated emergent learning of structural information, predicting residues involved in binding interactions.
- Outperformed single-omic models in both multi-omic and single-omic tasks on a performance-per-FLOP basis.
Conclusions:
- Multi-omic transformer models like OmniBioTE offer a powerful unified approach for biological sequence analysis.
- Integrating diverse omics data enhances predictive accuracy and reveals deeper biological insights.
- OmniBioTE sets a new standard for open-source tools in multi-omic bioinformatics.
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Organization
The primary structure of a protein is its amino acid sequence....

