XGBoost-增强组合模型使用歧视性混合特征来预测化位点
Salman Khan1, Sumaiya Noor2, Tahir Javed3
1New Emerging Technologies and 5G Network and Beyond Research Chair, Department of Computer Engineering, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia.
BioData mining
|February 3, 2025
概括
这项研究介绍了XGBoost-Sumo,这是一种用于预测化地点的新型计算模型. 这一进步有助于理解蛋白质功能和疾病,在识别这些关键的翻译后修改方面具有高度准确性.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 翻译后修饰 (PTMs) 调节关键细胞过程,包括基因表达和蛋白质稳定性.
- 合,一个关键的PTM,对蛋白质功能至关重要,并与帕金森氏症和阿尔茨海默氏症等神经退行性疾病有关.
- 准确识别化部位对于了解蛋白调节和疾病机制至关重要.
研究的目的:
- 开发一个强大而准确的计算模型来预测化场所.
- 整合各种蛋白质数据,包括结构和序列,以改善预测.
- 增强对化在生物功能和疾病中的作用的理解.
主要方法:
- 开发了XGBoost-Sumo,这是一个结合基于变压器的编码和PsePSSM-DWT用于进化特征提取的模型.
- 融合了词嵌入与进化描述符,并使用了SHapley添加式扩展 (SHAP) 进行特征选择.
- 使用极端梯度提升 (XGBoost) 来分类聚合位点.
主要成果:
- 通过十倍交叉验证,XGBoost-Sumo在基准数据集上实现了99.68%的准确性.
- 该模型在独立的测试样本上显示了96.08%的准确性.
- 在培训方面表现明显优于现有模型,在培训方面表现10.31%,在独立数据方面表现2.74%.
结论:
- XGBoost-Sumo提供了一种高度可靠和准确的方法来预测化地点.
- 该模型的性能代表了PTM预测领域的重大进步.
- 这种工具具有强大的潜力,可以加速制药开发和生物研究.
相关概念视频
Sensitivity, Specificity, and Predicted Value
171
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
171
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Hybridoma Technology
14.0K
Hybridoma technology is used for the large-scale production of monoclonal antibodies. Monoclonal antibodies bind to only a single antigenic determinant or epitope. Such antibodies are used in research, diagnostics, and disease therapy. The hybridoma technology established in 1975 by Georges Köhler and Cesar Milstein was awarded the Nobel Prize in Medicine in 1984 for revolutionizing research and therapy.
Hybridoma Selection
Commonly used fusion techniques — electroporation,...
Hybridoma Selection
Commonly used fusion techniques — electroporation,...
14.0K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Improving Translational Accuracy
2.5K
2.5K


