Related Experiment Video
Updated: Aug 13, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Synthetic sequence alignments as programmable probes of learned conformational landscapes in deep learning protein
Jannik Adrian Gut1,2, Noah Kleinschmidt1,2, Thomas Lemmin1
1Institute of Biochemistry and Molecular Medicine, University of Bern, Bühlstrasse 28, 3012, Bern, Switzerland.
Motivation:
Proteins rely on conformational flexibility for biological function, yet predicting alternative states remains a major challenge in structural biology. Although deep learning models like AlphaFold2, AlphaFold3, and RoseTTAFold2 excel at static structure prediction, what these networks actually learn about underlying conformational landscapes remains largely opaque.
Results:
Here we introduce synthetic multiple sequence alignments (MSAs), designed by inverse folding to encode predefined structural constraints, as a programmable intervention for interrogating the internal logic of structure prediction systems. Synthetic MSAs systematically bias AlphaFold2, AlphaFold3, and RoseTTAFold2 toward distinct conformational states of fold-switching proteins, including alternative conformations inaccessible through natural sequence information alone. Adversarial experiments pairing query sequences with MSAs encoding competing folds reveal sequence-dependent responses, exposing how alignment-derived and sequence-derived signals are weighted within each system. Probing predictions initialized from molecular dynamics trajectories reveals a systematic bias toward compact, training-distribution-favored conformations. Hybrid alignments combining synthetic and natural MSA segments enable targeted steering toward specific conformational states. These results suggest synthetic MSAs as a generalizable framework for dissecting the conformational landscapes encoded by deep learning structure predictors, with direct implications for understanding model behavior and accessing biologically relevant hidden states.
Availability And Implementation:
Newly generated data can be found on https://zenodo.org/records/20916910. The code underlying this article is available in GitHub at https://github.com/ibmm-unibe-ch/msa-tests.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein Organization
The primary structure of a protein is its amino acid sequence.
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme can...
Protein-protein Interfaces

