跨祖先信息传输框架改善了蛋白质丰度预测和蛋白质特征关联识别
Wenli Zhai1,2, Lingyun Sun1,2, Wenwei Fang1,2
1The Second Affiliated Hospital and School of Public Health, Zhejiang University School of Medicine, 866 Yuhangtang Road, Hangzhou, Zhejiang 310058, China.
Briefings in bioinformatics
|January 7, 2026
概括
一个新的多祖先最佳表现模型 (MABM) 通过增强代表性不足的人群中的蛋白质预测来改进全蛋白质组关联研究 (PWAS). 这种方法可以识别复杂疾病的更多遗传关联,有助于基因发现.
科学领域:
- 遗传学 是一个遗传学.
- 蛋白质组学是指蛋白质组学.
- 计算生物学 计算生物学
背景情况:
- 以遗传学为基础的全蛋白质组关联研究 (PWAS) 对于理解复杂的疾病机制至关重要.
- 目前的PWAS方法依赖于祖先匹配的参考面板,这些面板仅限于代表性不足的人群.
- 这种限制阻碍了在各种祖先中发现与疾病相关的蛋白质.
研究的目的:
- 开发一种新的多祖先框架,以提高代表性不足的人群中蛋白质预测的准确性.
- 增强在多种祖先中具有复杂特征的蛋白质组关联的识别.
- 为了促进基因和蛋白质优先考虑在多omics研究中的功能验证.
主要方法:
- 开发了一个多祖先最佳表现模型 (MABM),整合了各种信息共享策略.
- 应用MABM以提高蛋白质预测性能在交叉验证和外部数据集.
- 利用日本生物银行的PWAS数据集,并将MABM与拉索模型进行比较.
主要成果:
- 在不同的人群中,MABM显著提高了蛋白质预测性能.
- 与拉索模型相比,MABM在日本生物银行数据集中发现了PWAS关联的三倍.
- 47.5%的MABM特定关联在独立的东亚数据集中成功复制,显示出强度.
结论:
- MABM框架有效地提高了蛋白质预测和PWAS在代表性不足的人群中的发现.
- MABM促进了新型特征相关蛋白质候选者的识别,并验证了已知的关联.
- 这种方法扩大了多学科研究的适用性,特别是在代表性不足的群体中,并有助于发现与特征相关的蛋白质.
相关概念视频
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
Protein Networks
4.5K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.5K
Conservation of Protein Domains
3.9K
3.9K
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K


