Related Experiment Video
Updated: Aug 22, 2025

A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
End-to-end learning of multiple sequence alignments with differentiable Smith-Waterman
Samantha Petti1, Nicholas Bhattacharya2, Roshan Rao3
1NSF-Simons Center for the Mathematical and Statistical Analysis of Biology, Harvard University, Cambridge, MA 02138, USA.
We developed a differentiable alignment method to jointly learn multiple sequence alignments (MSAs) and downstream tasks. This approach improves protein structure prediction and contact prediction by creating novel, albeit adversarial, alignments.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning
Background:
- Multiple sequence alignments (MSAs) are crucial for understanding sequence-structure-function relationships and evolutionary histories.
- Current MSA generation is often a separate step, lacking integration with downstream applications like structure prediction.
Purpose of the Study:
- To develop a method for jointly learning MSAs and downstream machine learning systems in an end-to-end manner.
- To introduce SMURF (Smooth Markov Unaligned Random Field) for unsupervised contact prediction using learned alignments.
- To demonstrate the utility of differentiable alignment modules in improving protein structure prediction pipelines.
Main Methods:
- Implemented a smooth and differentiable version of the Smith-Waterman algorithm for joint MSA learning.
- Developed SMURF, integrating differentiable alignment with Markov Random Fields for contact prediction.
- Connected the differentiable alignment module to AlphaFold2 to optimize MSAs for structure prediction.
Main Results:
- SMURF achieved mild improvements in contact prediction for protein and RNA families.
- Jointly learned MSAs improved AlphaFold2 structure predictions compared to initial MSAs.
- Alignments that enhanced structure prediction were found to be self-inconsistent, suggesting adversarial properties.
Conclusions:
- Differentiable dynamic programming offers a powerful approach to enhance neural network pipelines reliant on sequence alignments.
- Optimizing protein sequence predictions using current methods requires careful consideration due to potential adversarial effects.
- The developed method facilitates end-to-end learning for MSA-dependent biological applications.
More Related Videos
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Evolutionary Relationships through Genome Comparisons
Conservation of Protein Domains
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....