Related Experiment Video
Updated: May 3, 2026

09:22
Multi-Faceted Mass Spectrometric Investigation of Neuropeptides in Callinectes sapidus
Published on: May 31, 2022
2.0K
Signal peptide discrimination and cleavage site identification using SVM and NN
H B Kazemian1, S A Yusuf2, K White1
1London Metropolitan University, UK.
Computers in Biology and Medicine
|February 1, 2014
Summary
This study introduces a novel Support Vector Machine (SVM)-Neural Network (NN) method to accurately identify signal peptide (SP) sequences and their cleavage sites in proteins, improving protein topology modeling.
Area of Science:
- Bioinformatics
- Computational Biology
- Proteomics
Background:
- Signal peptides (SPs) target proteins to secretory pathways, with ~15% of genomic proteins possessing them.
- Accurate SP prediction is vital for membrane protein topology modeling, as SPs can be mistaken for transmembrane domains.
- The cleavage of SPs releases mature proteins, a critical step in protein processing.
Purpose of the Study:
- To develop and validate a cascaded SVM-NN classification methodology for discriminating signal peptide (SP) sequences from non-SP sequences.
- To accurately identify the cleavage sites of SPs using a two-phase classification approach.
- To enhance the prediction accuracy for protein targeting and processing.
Main Methods:
- A dual-phase classification approach was employed, utilizing SVM for primary SP/Non-SP discrimination and NN for cleavage site prediction.
- Phase one involved SVM classification with hydrophobic propensities and symmetric sliding windows for SP discrimination.
- Phase two used NN classification with asymmetric sliding windows for precise cleavage site identification.
Main Results:
- The SVM-NN model achieved an overall accuracy of 0.90 for SP and Non-SP discrimination, validated by Matthews Correlation Coefficient (MCC) tests.
- The methodology demonstrated a high accuracy of 91.5% for SP cleavage site prediction through cross-validation.
- The proposed method was tested on Uni-Prot non-redundant datasets of eukaryotic and prokaryotic proteins.
Conclusions:
- The cascaded SVM-NN methodology provides a robust and accurate approach for signal peptide identification and cleavage site prediction.
- This method significantly improves the accuracy of modeling protein topology and understanding protein secretion pathways.
- The findings offer a valuable tool for bioinformatics and computational biology research, particularly in proteomics.
Related Concept Videos
Signal Sequences and Sorting Receptors
9.9K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
9.9K
Peptide Identification Using Tandem Mass Spectrometry
6.2K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.2K
Nuclear Localization Signals and Import
6.2K
Proteins targeted to the nucleus carry short stretches of amino acid sequences called the nuclear localization signal or NLS. Classical nuclear localization signals are of two types: monopartite and bipartite NLS. Monopartite classical NLS (cNLS) consists of a single cluster of 4-8 amino acids. Bipartite cNLS consists of two clusters of 2-3 amino acids and a 9-12 residue long proline-rich linker bridging the two clusters. Signal clusters are rich in positively charged amino acids such as...
6.2K

