序列聚类混了AlphaFold2的情况
Joseph W Schafer1, Devlina Chakravarty1, Ethan A Chen1
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894.
bioRxiv : the preprint server for biology
|February 5, 2024
概括
AF集群努力预测变形蛋白质结构,经常将单折叠的蛋白质误认为折叠切换器. 使用ColabFold随机序列采样为预测蛋白质结构提供了更可靠和更有效的替代方案.
科学领域:
- 蛋白质结构预测 蛋白质结构预测
- 计算生物学是一种计算生物学.
- 生物物理学的生物物理.
背景情况:
- 球状蛋白质可以在多个构造之间切换,称为变形蛋白质.
- AlphaFold2 (AF2) 准确地预测了主要的蛋白质结构,但往往无法捕捉出替代性构造.
- 预测这些替代结构对于理解蛋白质的功能和调节至关重要.
研究的目的:
- 为了评估AF-cluster的可靠性,该方法旨在使用AlphaFold2.2.来预测变形蛋白质的替代结构.
- 确定AF集群方法中的局限性和潜在偏差.
- 提出替代的,更准确的方法来预测变形蛋白质结构.
主要方法:
- 分析已公布的AF集群结果,并与随机序列采样进行比较.
- 对已知的单折叠蛋白 (KaiB同类) 的AF集群性能评估.
- 预测结构信心得分与实验观测的错误分析.
主要成果:
- 随机序列采样与AF集群的序列集群相比,显示出更高的性能.
- AF集群错误地将单折叠的KaiB同类分子识别为折叠切换蛋白.
- AF集群在预测正确结构方面表现出低信心,在预测未观察到的形状方面表现出高信心.
结论:
- 由于方法上的缺陷,AF-集群是变形蛋白质结构的不可靠预测器.
- 该方法错误地分类了单折叠蛋白质,并表现出不良的信心校准.
- 推基于ColabFold的随机序列采样,可能与其他方法相辅助,作为更准确和计算效率更高的替代方案.
相关概念视频
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K
Protein Folding
118.2K
Overview
118.2K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


