准确预测抗癌,使用一个堆叠组合的卷积和变压器模型与联合序列表示
Anh Duy Huynh1, Phurinut Khampasri2, Pimmada Janthanet2
1Graduate School in the Program of Research and Development in Pharmaceuticals, Faculty of Pharmaceutical Sciences, Khon Kaen University, Khon Kaen, 40002, Thailand; Department of Health Sciences, College of Natural Sciences, Can Tho University, Can Tho, 900000, Viet Nam.
Computers in biology and medicine
|January 15, 2026
概括
我们使用先进的机器学习开发了一种强大的计算工具,用于识别抗癌 (ACP). 该模型准确地预测了非洲和非洲国家和地区的情况,并有助于发现新的基于的癌症治疗方法.
科学领域:
- 生物技术是生物技术.
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 抗癌 (ACP) 在癌症治疗方面表现有前途.
- 准确识别ACP对于药物发现至关重要.
- 现有的计算方法在性能和可解释性方面存在局限性.
研究的目的:
- 开发一个高性能,可解释和可概括的预测框架,用于ACP国家和地区的识别.
- 整合多样化的序列表示,以进行全面的特征提取.
- 创建一个用户友好的工具,以促进ACP的发现.
主要方法:
- 一种堆叠集体学习方法,结合了卷积神经网络和变压器模型.
- 联合序列表示集成一热编码和进化规模建模嵌入.
- 为了模型的可解释性,夏普利添加式解释 (SHAP).
主要成果:
- 在主要的非洲和非洲国家和地区数据集上获得了88.9%的准确度,外部验证准确度从83.2%到95.2%不等.
- 具有强大的概括能力,与最先进的模型相比.
- 确定了像氨酸和其他充电/疏水氨基酸这样的关键残留物,这些残留物对非洲和非洲的活动有影响.
结论:
- 拟议的框架为计算ACP预测提供了一个强大的和可解释的方法.
- 这些发现支持了对ACP-膜相互作用的机制性见解,并指导了的设计.
- 一个交互式的网络工具,ACPredictor,可用于提高可访问性和加速基于的抗癌药物发现.
相关概念视频
Peptide Identification Using Tandem Mass Spectrometry
8.1K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
8.1K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K

