通过AlphaFold2系统识别有条件折叠的内在无序区域
T Reid Alderson1,2, Iva Pritišanac3,4,5, Đesika Kolarić5
1Department of Biochemistry, University of Toronto, Toronto, ON M5S 1A8, Canada.
概括
AlphaFold2令人惊地预测了有条件折叠的内在无序区域 (IDR) 的结构. 这些区域与人类疾病有关,并且在原核生物中很丰富,但不是在真核生物中.
科学领域:
- 结构生物学是结构生物学.
- 生物信息学是一种生物信息学.
- 计算生物学是一种计算生物学.
背景情况:
- 阿尔法蛋白质结构数据库提供了数以百万计的预测蛋白质结构.
- 本质上无序的区域 (IDR) 缺乏稳定的结构,通常被认为具有较低的预测信心.
- 条件折叠描述了在特定条件 (如绑定) 上采用结构的IDR.
研究的目的:
- 调查AlphaFold2对内在无序区域 (IDR) 的信任分数.
- 评估AlphaFold2能够预测条件折叠IDR的结构的能力.
- 探索条件折叠IDRs,疾病突变和家族遗传分布之间的关系.
主要方法:
- 对人类内在无序区域 (IDR) 的 AlphaFold2 置信度评分的分析.
- 将AlphaFold2预测与条件折叠IDR实验核磁共振 (NMR) 数据进行比较.
- 使用已知条件折叠IDRs的数据库评估AlphaFold2的性能.
- 在条件折叠的IDR中对人类疾病突变丰富的分析.
- 在 prokaryotic 和 eukaryotic IDR 中预测条件折叠的比较分析.
主要成果:
- AlphaFold2将可靠的结构预测赋予了近15%的人类内在无序区域 (IDR).
- AlphaFold2准确地预测了NMR数据验证的条件折叠IDR子集的折叠状态结构.
- 鉴定条件折叠IDR的估计精度高达88%,假阳性率为10%.
- 人类疾病突变在有条件折叠的IDR中显著丰富 (几乎是五倍).
- 预计高达80%的 prokaryotic IDRs 会有条件折叠,相比之下,只有不到20%的真核IDRs.
结论:
- 尽管训练数据有限,但AlphaFold2能够非常准确地识别有条件折叠的内在无序区域 (IDR).
- 有条件折叠的IDR是人类疾病突变的热点.
- 条件折叠在 prokaryotes 中普遍存在,但在 eukaryotes 中不太常见,这表明不同的功能角色.
- AlphaFold2的预测不能完全捕捉到IDR的功能可塑性或组合性质.
相关概念视频
Protein Folding
118.3K
Overview
118.3K
Protein Folding Quality Check in the RER
3.7K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.7K
Molecular Chaperones and Protein Folding
18.0K
The native conformation of a protein is formed by interactions between the side chains of its constituent amino acids. When the amino acids cannot form these interactions, the protein cannot fold by itself and needs chaperones. Notably, chaperones do not relay any additional information required for the folding of polypeptides; the native conformation of a protein is determined solely by its amino acid sequence. Chaperones catalyze protein folding without being a part of the folded protein.
The...
The...
18.0K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Intrinsically Disordered Proteins
17.9K
Intrinsically disordered proteins are a group of proteins that do not fold into specific three-dimensional structures. Their structural flexibility allows them to complement ordered proteins to perform functions that are inaccessible to rigid structures. They are more common in eukaryotes than prokaryotes and may either be exclusively intrinsically disordered or hybrid proteins, consisting of a mix of ordered and disordered regions. The absence of a rigid structure in these proteins can be...
17.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


