Related Experiment Video
Updated: Jun 13, 2026

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
Phylogenetic corrections and higher-order sequence statistics in protein families: Potts vs multiple sequence
Kisan Khatri1, Ronald M Levy2, Allan Haldane1
1Department of Physics and Center for Biophysics and Computational Biology, Temple University, Philadelphia, Pennsylvania 19122, USA.
Abstract:
Recent generative machine learning models applied to protein multiple sequence alignment (MSA) datasets include simple and interpretable physics-based Potts covariation models and other machine learning models such as MSA Transformer (MSA-T). The best models accurately reproduce MSA statistics induced by the biophysical constraints within proteins, raising the question of which functional forms best model the underlying physics. The Potts model is usually specified by an effective potential including pairwise residue-residue interaction terms, but it has been suggested that MSA-T can capture the effects induced by effective potentials that include more than pairwise interactions and implicitly account for phylogenetic structure in the MSA. Here, we compare the ability of the Potts model and MSA-T to reconstruct higher-order sequence statistics reflecting complex biophysical sequence constraints. We show how the model performance depends greatly on the treatment of phylogenetic relationships between the sequences, which can induce nonbiophysical mutational covariation in MSAs. When using explicit corrections for phylogenetic dependencies, we find that the Potts model outperforms MSA-T in detecting epistatic interactions of biophysical origin. Furthermore, the sequences generated by the models we studied here are predicted to fold into nativelike structures, as indicated by AlphaFold predictions.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Protein Families
Protein Families
Microbial Phylogeny
Phylogenetic Trees
Phylogenetic Trees

