Related Experiment Video
Updated: Oct 5, 2025

06:19
Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
985
Positional SHAP (PoSHAP) for Interpretation of machine learning models trained from biological sequences.
Quinn Dickinson1, Jesse G Meyer1
1Department of Biochemistry, Medical College of Wisconsin, Milwaukee, Wisconsin.
Plos Computational Biology
|January 28, 2022
Summary
Positional SHAP (PoSHAP) offers a new way to interpret deep learning models for biological sequences. This framework helps understand predictions for peptide properties like MHC binding and collisional cross section.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning
Background:
- Deep learning models, particularly recurrent neural networks, excel at biological predictions from sequential data.
- Interpreting these complex models, especially for sequential inputs, remains a significant challenge.
Purpose of the Study:
- To introduce Positional SHAP (PoSHAP), a novel framework for interpreting machine learning models trained on biological sequences.
- To generate positional model interpretations using SHapely Additive exPlanations (SHAP).
Main Methods:
- Developed and applied the Positional SHAP (PoSHAP) framework.
- Utilized SHapely Additive exPlanations (SHAP) to derive positional interpretations.
- Tested PoSHAP on three long short-term memory (LSTM) regression models predicting peptide properties (MHC binding affinity, collisional cross section).
Main Results:
- PoSHAP successfully reproduced known peptide binding motifs for MHC class I molecules (Mamu-A1*001 and A*11:01).
- The framework accurately reflected established properties of peptide collisional cross section (CCS).
- New insights into interpositional dependencies of amino acid interactions within peptide sequences were uncovered.
Conclusions:
- Positional SHAP (PoSHAP) provides effective positional interpretations for models trained on biological sequences.
- The framework demonstrates broad utility for understanding various sequence-based biological prediction models.
- PoSHAP aids in deciphering complex relationships within biological sequence data.
More Related Videos
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
6.4K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.4K
Sequence Networks of Rotating Machines
165
A Y-connected synchronous generator, grounded through a neutral impedance, is designed to produce balanced internal phase voltages with only positive-sequence components. The generator's sequence networks include a source voltage that is exclusively in the positive-sequence network. The sequence components of line-to-ground voltages at the generator terminals illustrate this configuration.
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
165
Conserved Binding Sites
4.5K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.5K
Signal Sequences and Sorting Receptors
9.1K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
9.1K
Phylogenetic Trees
47.9K
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.
47.9K
Genome Annotation and Assembly
19.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.5K

