西格玛利用蛋白质结构信息来预测误解变异的病原性
Hengqiang Zhao1, Huakang Du1, Sen Zhao1
1Department of Orthopedic Surgery, State Key Laboratory of Complex Severe and Rare Diseases, Peking Union Medical College Hospital, Peking Union Medical College and Chinese Academy of Medical Sciences, Beijing 100730, China; Beijing Key Laboratory for Genetic Research of Skeletal Deformity, Beijing 100730, China.
Cell reports methods
|January 11, 2024
概括
我们开发了SIGMA,这是一种使用蛋白质结构预测来评估误解变体致病性的新工具. 在SIGMA+中与其他预测器相结合,SIGMA在识别引起疾病的基因突变方面表现出卓越的准确性.
科学领域:
- 基因组学就是基因组学.
- 结构生物学 结构生物学
- 生物信息学是一种生物信息学.
背景情况:
- 评估误解变异的致病性对于遗传疾病诊断至关重要.
- 有限的实验确定蛋白质结构阻碍了病原性预测.
- 需要计算方法来克服结构数据的稀缺性.
研究的目的:
- 开发一种结构信息化工具,用于预测误解变体的病原性.
- 为了利用AlphaFold2预测进行增强的致病性评估.
- 改进现有的误解变异病原性预测指标.
主要方法:
- 开发了使用AlphaFold2蛋白质结构预测的结构信息基因误解突变评估器 (SIGMA).
- 评估了SIGMA在标记变体和实验数据集上的性能.
- 将SIGMA与其他预测器集成,以创建SIGMA+.
- 在疾病相关基因中计算了数百万个变异的SIGMA分数.
主要成果:
- 在预测误解变异病原性 (AUC = 0.933) 方面,SIGMA表现优越.
- 突变残留物的相对溶剂可访问性对SIGMA的预测能力做出了重大贡献.
- 通过结合预测因素,SIGMA+实现了更高的准确性 (AUC = 0.966).
- 为了广泛应用SIGMA,推出了一个交互式在线平台.
结论:
- SIGMA提供了一个准确的,基于结构的方法,用于误解变异病原性评估.
- 利用预测的蛋白质结构提高了病原性预测能力.
- 西格玛和西格玛+为遗传变异解释和研究提供了宝贵的资源.
相关概念视频
Structural Protein Function
2.8K
2.8K
Single Nucleotide Polymorphisms-SNPs
15.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K
Signal Sequences and Sorting Receptors
5.4K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.4K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Protein Folding Quality Check in the RER
3.7K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.7K
Mutations
82.4K
Overview
82.4K


