Computational prediction resolves thousands of homooligomeric phage protein structures.
Susanna R Grigson1,2, Natia Geliashvili3,4, Torsten Schubert3,4
1Flinders Accelerator for Microbiome Exploration, College of Science and Engineering, Flinders University, Adelaide, SA, 5042, Australia.
Biorxiv : the Preprint Server for Biology
|June 5, 2026
Summary
We developed PHLEGM, a computational method to predict phage homooligomeric states from protein sequences. This tool accurately identifies protein complexes, aiding the study of viral diversity and function.
Area of Science:
- Structural biology
- Computational biology
- Virology
Background:
- Bacteriophages (phages) are crucial in microbial ecosystems, but their proteins are largely uncharacterized.
- Understanding protein structure, especially quaternary structure, is key to determining protein function.
- Many phage proteins form homooligomers (complexes of identical subunits), making their prediction a significant interest.
Purpose of the Study:
- To develop and validate a computational framework, PHLEGM, for predicting homooligomeric states of phage proteins directly from their amino acid sequences.
- To assess the accuracy and utility of PHLEGM by experimental validation and benchmarking against other methods.
- To explore the diversity of phage homooligomeric protein complexes within a large phage protein database.
Main Methods:
- Developed PHLEGM, integrating AlphaFold-Multimer modeling with interface quality assessment for homooligomer prediction from protein sequences.
- Experimentally validated two predicted homooligomers (a dimer and a trimer) using size exclusion chromatography and hydrodynamic techniques.
- Applied PHLEGM to over 22,000 phage protein sequences from the PHROGs database and benchmarked against protein language model predictors.
Main Results:
- PHLEGM successfully predicted homooligomeric states, with experimental validation confirming predictions for a dimer and a trimer.
- Application to the PHROGs database revealed significant diversity in phage homooligomeric protein complexes.
- PHLEGM demonstrated superior accuracy compared to protein language model-based predictors, particularly for trimers and higher-order complexes.
Conclusions:
- PHLEGM is a robust, structure-based computational tool for predicting phage homooligomeric states, advancing the understanding of viral protein function.
- The study highlights the value of experimentally benchmarked computational predictions in deciphering the vast phage sequence space.
- Publicly released predicted structures and functional inferences will facilitate future research on phage proteins.
Related Concept Videos
DNA Bacteriophages
Bacteriophages, or phages, are viruses that specifically infect bacteria, utilizing their genetic material to hijack host cellular machinery for replication. DNA bacteriophages employ single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA) genomes. These phages exhibit diverse replication strategies and host interactions, influencing their ecological roles and applications in biotechnology and medicine.ssDNA BacteriophagesssDNA phages, with their small genomes, utilize unique strategies to...
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.
Protein Complex Assembly
Proteins can form homomeric complexes with another unit of the same protein or heteromeric complexes with different types. Most protein complexes self-assemble spontaneously via ordered pathways, while some proteins need assembly factors that guide their proper assembly. Despite the crowded intracellular environment, proteins usually interact with their correct partners and form functional complexes.
Many viruses self-assemble into a fully functional unit using the infected host cell to...
Many viruses self-assemble into a fully functional unit using the infected host cell to...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...


