乌比戈-X:使用组合学习与基于图像的特征表示和加权投票的蛋白质无处不在部位预测
Disline Manli Tantoh1, Jen-Chieh Yu2, Ching-Hsuan Chien1
1Doctoral Program in Medical Biotechnology, National Chung Hsing University, Taichung City, Taiwan.
Computational and structural biotechnology journal
|July 29, 2025
概括
一个新的工具Ubigo-X通过整合序列和结构特征,准确地预测蛋白质无处不在的位置. 它的性能优于现有的方法,为生物功能分析提供了宝贵的资源.
科学领域:
- 生物化学和分子生物学
- 生物信息学和计算生物学
- 蛋白质组学是指蛋白质组学.
背景情况:
- 准确识别无处不在的部位对于理解生物功能至关重要.
- 现有的预测工具可能缺乏足够的准确性或通用性.
研究的目的:
- 开发一种新的,准确的蛋白质无处不在预测工具,名为Ubigo-X.
- 整合各种特征类型,包括序列,结构和功能,以提高预测性能.
- 提供一种物种中立的工具来预测无处不在的地点.
主要方法:
- 开发了三个子模型:单类型基于序列的特征 (SBF),k-mer基于序列的特征 (Co-Type SBF) 和基于结构/功能的特征 (S-FBF).
- 使用的氨基酸成分,AAindex,一次热编码,k-mer编码,二次结构,溶剂可访问性和信号裂部位.
- 采用S-FBF的XGBoost和基于图像的SBF和Co-Type SBF的Resnet34,通过加权投票组合模型.
主要成果:
- 在独立数据集上,Ubigo-X实现了高性能:在平衡的PhosphoSitePlus数据上,AUC为0.85,ACC为0.79,MCC为0.58.
- 在不平衡数据 (AUC 0.94,ACC 0.85,MCC 0.55) 和GPS-Uber数据 (AUC 0.81,ACC 0.59,MCC 0.27) 上表现出强大的性能.
- 在平衡数据的马修斯相关系数 (MCC) 和准确性 (ACC) 中表现优于现有的工具.
结论:
- Ubigo-X有效地整合了各种功能,使用基于图像的表示和加权的投票进行了卓越的无处不在预测.
- 该工具显示出作为物种中立预测器的巨大潜力,用于无处不在的地点.
- Ubigo-X可以在线访问,为该领域的研究人员提供了宝贵的资源.
相关概念视频
Peptide Identification Using Tandem Mass Spectrometry
6.8K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.8K
The Proteasome
1.2K
Eukaryotic cells can degrade proteins through several pathways. One of the most important among these is the ubiquitin-proteasome pathway. It helps the cell eliminate the misfolded, damaged, or unwarranted cytoplasmic proteins in a highly specific manner.
In this pathway, the target proteins are first tagged with small proteins called ubiquitin. This involves participation of a series of enzymes including— E1 (ubiquitin-activating enzyme), E2 (ubiquitin-conjugating enzyme), and E3...
In this pathway, the target proteins are first tagged with small proteins called ubiquitin. This involves participation of a series of enzymes including— E1 (ubiquitin-activating enzyme), E2 (ubiquitin-conjugating enzyme), and E3...
1.2K
Weighted Mean
5.3K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
5.3K
Conservation of Protein Domains Over Different Proteins
11.4K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.4K
Conserved Binding Sites
4.4K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.4K
Proteomics
7.9K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.9K


