Related Experiment Video
Updated: Feb 3, 2026

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
69.8K
Single-sequence-based prediction of protein secondary structures and solvent accessibility by deep whole-sequence
Rhys Heffernan1, Kuldip Paliwal1, James Lyons1
1Signal Processing Laboratory, Griffith University, Brisbane, QLD, 4111, Australia.
Journal of Computational Chemistry
|October 29, 2018
Summary
This study introduces SPIDER3-Single, a novel single-sequence protein structure prediction method. It achieves high accuracy without evolutionary data, offering an efficient alternative for analyzing protein structural properties.
Area of Science:
- Computational biology
- Structural bioinformatics
- Machine learning in protein science
Background:
- Protein structure prediction is crucial for understanding protein function.
- Traditional methods heavily rely on evolutionary information from multiple sequence alignments.
- Long Short-Term Bidirectional Recurrent Neural Networks (LSTM-BRNNs) show promise in capturing sequence dependencies.
Purpose of the Study:
- To develop and evaluate a single-sequence-based protein structure prediction method.
- To assess the performance of LSTM-BRNNs for predicting various protein structural properties.
- To provide an efficient computational tool for large-scale protein analysis.
Main Methods:
- Utilized Long Short-Term Bidirectional Recurrent Neural Networks (LSTM-BRNNs) for prediction.
- Developed a method named SPIDER3-Single, operating solely on single protein sequences.
- Evaluated prediction accuracy for Q3, solvent accessible surface area, secondary structure, and backbone angles.
Main Results:
- SPIDER3-Single achieved a Q3 accuracy of 72.5% and a 0.67 correlation coefficient for solvent accessible surface area.
- The method accurately predicted eight-state secondary structure, main-chain angles, half-sphere exposure, and contact number.
- Outperformed evolutionary-based methods for proteins with limited sequence homologs.
Conclusions:
- Single-sequence-based prediction using LSTM-BRNNs is a viable and accurate approach.
- SPIDER3-Single offers a computationally efficient alternative for protein structure property prediction.
- The method is accessible via the SPIDER3 server and as a standalone download.
Keywords:
backbone anglescontact predictionprotein structure predictionsecondary structure predictionsolvent accessibility predictionMore Related Videos
Related Concept Videos
Cis-regulatory Sequences
11.8K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.8K
Cis-regulatory Sequences
4.1K
4.1K
Protein and Protein Structure
87.6K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
87.6K
Sequences
277
Sequences are fundamental mathematical objects consisting of ordered lists of numbers that follow a specific rule or pattern. Sequences are critical in various mathematical concepts, including calculus, series, and number theory. They can model real-world phenomena such as population growth, financial investments, and physical processes like the diminishing height of a bouncing ball.Each number in a sequence is referred to as a term. Typically, the terms are denoted as a1, a2, a3,…, where...
277
Sanger Sequencing
774.5K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
774.5K
Arithmetic Sequences
239
An arithmetic sequence is a structured arrangement of numbers where each term is derived by adding a constant value, known as the common difference, to the previous term. This consistent pattern allows for the efficient computation of any term within the sequence as well as the cumulative sum of multiple terms. The formula for finding the nth term of an arithmetic sequence is:Here, aₙ represents the nth term of the sequence, a is the first term, d is the common difference, and n is the...
239

