基于改进的FCN和双向LSTM的谷物蛋白功能预测
Jing Liu1, Kun Li1, Xinghua Tang1
1College of Information Engineering, Shanghai Maritime University, Shanghai 201306, China.
Food chemistry
|April 10, 2025
概括
一个新的PBiLSTM-FCN模型通过考虑氨基酸序列顺序和长期依赖,准确地预测谷物蛋白的功能. 这种生物信息学方法提高了对大豆和玉米等作物的蛋白质功能预测准确度.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 高通量测序需要先进的方法来从氨基酸序列预测蛋白质功能.
- 现有的模型往往忽略了氨基酸链中的关键序列顺序和远程依赖关系.
- 准确的蛋白质功能预测对于理解和改善作物特征至关重要.
研究的目的:
- 开发和评估一种用于预测谷物蛋白功能的新型智能模型.
- 为了解决捕获氨基酸序列顺序和长期依赖性的现有方法的局限性.
- 提高重要的作物物种中蛋白质功能预测的准确性和可靠性.
主要方法:
- 利用了数据集,其中包括来自UniProtKB.来源的大豆,玉米,印加和日本豆的谷物蛋白.
- 提出了PBiLSTM-FCN模型,集成完全卷积网络 (FCN) 进行序列顺序和双向长期短期内存 (BiLSTM) 进行长期依赖.
- 对现有模型进行实验性比较,以评估性能.
主要成果:
- 与现有方法相比,PBiLSTM-FCN模型显示出更高的性能.
- 该模型有效地捕获了长距离的依赖关系和氨基酸序列的顺序,从而提高了预测准确度.
- 解释性分析通过将预测的功能与实际的蛋白质功能进行比较,证实了该模型的有效性.
结论:
- PBiLSTM-FCN模型在预测谷物蛋白功能方面取得了重大进展.
- 该模型处理序列顺序和长期依赖性的能力为生物信息学任务提供了更准确的方法.
- 这项工作为功能基因组学和作物改进研究提供了宝贵的工具.
相关概念视频
Protein Complexes with Interchangeable Parts
2.5K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.5K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Genome Annotation and Assembly
18.7K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.7K
lncRNA - Long Non-coding RNAs
8.4K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.4K
Protein Families
15.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.2K
Protein-protein Interfaces
12.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.4K


