Related Experiment Video
Updated: May 15, 2025

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
68.4K
Grain protein function prediction based on improved FCN and bidirectional LSTM
Jing Liu1, Kun Li1, Xinghua Tang1
1College of Information Engineering, Shanghai Maritime University, Shanghai 201306, China.
Food Chemistry
|April 10, 2025
Summary
A new PBiLSTM-FCN model accurately predicts grain protein function by considering amino acid sequence order and long-term dependencies. This bioinformatics approach enhances protein function prediction accuracy for crops like soybean and maize.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- High-throughput sequencing necessitates advanced methods for predicting protein function from amino acid sequences.
- Existing models often overlook the critical sequential order and long-range dependencies within amino acid chains.
- Accurate protein function prediction is vital for understanding and improving crop traits.
Purpose of the Study:
- To develop and evaluate a novel intelligent model for predicting grain protein function.
- To address the limitations of existing methods in capturing amino acid sequence order and long-term dependencies.
- To improve the accuracy and reliability of protein function prediction in important crop species.
Main Methods:
- Utilized a dataset comprising grain proteins from soybean, maize, indica, and japonica sourced from UniProtKB.
- Proposed the PBiLSTM-FCN model, integrating Fully Convolutional Networks (FCN) for sequence order and bidirectional Long Short-Term Memory (BiLSTM) for long-term dependencies.
- Conducted experimental comparisons against existing models to assess performance.
Main Results:
- The PBiLSTM-FCN model demonstrated superior performance compared to existing methods.
- The model effectively captured long-range dependencies and the order of amino acid sequences, leading to higher prediction accuracy.
- Interpretability analyses confirmed the model's effectiveness by comparing predicted functions with actual protein functions.
Conclusions:
- The PBiLSTM-FCN model represents a significant advancement in predicting grain protein function.
- The model's ability to handle sequence order and long-term dependencies offers a more accurate approach to bioinformatics tasks.
- This work provides a valuable tool for functional genomics and crop improvement research.
Related Concept Videos
Protein Complexes with Interchangeable Parts
2.5K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.5K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Genome Annotation and Assembly
18.7K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.7K
lncRNA - Long Non-coding RNAs
8.4K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.4K
Protein Families
15.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.2K
Protein-protein Interfaces
12.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.4K

