蛋白质:一种可解释的AI管道用于蛋白质分类,以支持精准医学
IEEE journal of biomedical and health informatics
|March 13, 2026
概括
我们开发了PROTEAN,这是一种人工智能工具,可以使用可解释的AI准确地分类可溶性N-乙基-马莱胺敏感因子附着蛋白受体 (SNARE) 蛋白质. 这提高了生物医学应用的透明度,例如疾病分析和生物标志物发现.
科学领域:
- 生物化学和生物信息学
- 人工智能在医学中的应用
- 计算生物学 计算生物学
背景情况:
- 溶性N-乙基-马利胺敏感因子附着蛋白受体 (SNARE) 蛋白质对于细胞内贩运和疾病至关重要.
- 由于序列和结构相似性,区分SNARE和NONSNARE蛋白质是很困难的.
- 传统的AI模型缺乏透明度,限制了它们在敏感的生物医学领域的使用.
研究的目的:
- 开发PROTEAN,一种结合机器学习和可解释AI (XAI) 的方法,用于准确和可解释的蛋白质分类.
- 为了应对区分SNARE和NONSNARE蛋白质的挑战.
- 提高AI在生物医学应用中的可靠性.
主要方法:
- 一个三个阶段的管道:在平衡数据集上进行数据预处理 (D128),对分类器 (SVM,NN) 进行培训和评估,并使用SHAP和LIME XAI模型进行解释.
- 使用了128个SNARE和NONSNARE蛋白序列的平衡数据集.
- 采用SHAP和LIME用于模型解释以确定关键蛋白质描述符.
主要成果:
- 中高斯支持向量机 (SVM) 分类器在458种蛋白质的测试组中实现了92.1%的准确性,94.8%的灵敏性和89.5%的特异性.
- SHAP和LIME为分类决定提供了一致的,生物相关的解释.
- 确定了氨基酸组成和序列顺序作为蛋白质分类的关键特征.
结论:
- PROTEAN表明,将可解释的AI与平衡的数据集集集成,可以提高蛋白质分类性能和透明度.
- 这些发现对于开发可靠的生物医学人工智能系统至关重要.
- 为疾病机制分析,生物标志物发现和个性化医疗提供了新的工具.
相关概念视频
Protein Networks
4.6K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.6K
Improving Translational Accuracy
15.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.4K
Improving Translational Accuracy
3.7K
3.7K


