Related Experiment Video
Updated: May 12, 2026

Synthesis and Structure Determination of µ-Conotoxin PIIIA Isomers with Different Disulfide Connectivities
Published on: October 2, 2018
On the relevance of sophisticated structural annotations for disulfide connectivity pattern prediction
Julien Becker1, Francis Maes, Louis Wehenkel
1Bioinformatics and Modeling, GIGA-Research, Department of Electrical Engineering and Computer Science, Montefiore Institute, University of Liege, Liege, Belgium.
Abstract:
Disulfide bridges strongly constrain the native structure of many proteins and predicting their formation is therefore a key sub-problem of protein structure and function inference. Most recently proposed approaches for this prediction problem adopt the following pipeline: first they enrich the primary sequence with structural annotations, second they apply a binary classifier to each candidate pair of cysteines to predict disulfide bonding probabilities and finally, they use a maximum weight graph matching algorithm to derive the predicted disulfide connectivity pattern of a protein. In this paper, we adopt this three step pipeline and propose an extensive study of the relevance of various structural annotations and feature encodings. In particular, we consider five kinds of structural annotations, among which three are novel in the context of disulfide bridge prediction. So as to be usable by machine learning algorithms, these annotations must be encoded into features. For this purpose, we propose four different feature encodings based on local windows and on different kinds of histograms. The combination of structural annotations with these possible encodings leads to a large number of possible feature functions. In order to identify a minimal subset of relevant feature functions among those, we propose an efficient and interpretable feature function selection scheme, designed so as to avoid any form of overfitting. We apply this scheme on top of three supervised learning algorithms: k-nearest neighbors, support vector machines and extremely randomized trees. Our results indicate that the use of only the PSSM (position-specific scoring matrix) together with the CSP (cysteine separation profile) are sufficient to construct a high performance disulfide pattern predictor and that extremely randomized trees reach a disulfide pattern prediction accuracy of [Formula: see text] on the benchmark dataset SPX[Formula: see text], which corresponds to [Formula: see text] improvement over the state of the art. A web-application is available at http://m24.giga.ulg.ac.be:81/x3CysBridges.
More Related Videos
09:37Combining Non-reducing SDS-PAGE Analysis and Chemical Crosslinking to Detect Multimeric Complexes Stabilized by Disulfide Linkages in Mammalian Cells in Culture
Published on: May 2, 2019
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence.
Structure and Nomenclature of Thiols and Sulfides
Protein Folding
Protein Modifications in the RER
Broadly, these modifications can be categorized into four main categories — glycosylation, formation of disulfide bonds, assembly of protein subunits, and specific proteolytic cleavages like removal of signal sequences.
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...