HapScoreDB:一个蛋白质语言模型功能评分的数据库,用于哈普类型解析的蛋白质序列
Fabio Mazza1, Filippo Gastaldello1,2, Davide Dalfovo1
1Department of Cellular, Computational and Integrative Biology (CIBIO), University of Trento, Trento 38123, Italy.
Nucleic acids research
|November 20, 2025
概括
HapScoreDB是一个新的数据库,使用蛋白质语言模型来预测人类单元型内的遗传变异的功能影响. 它揭示了癌症GWAS变体的单元类型的适应性降低,有助于基因解释.
科学领域:
- 人类遗传学 人类遗传学
- 计算生物学是一种计算生物学.
- 基因组学就是基因组学.
背景情况:
- 解释遗传变异的功能影响,特别是那些在单元型上共同遗传的变异,由于表观症而具有挑战性.
- 现有的数据库往往缺乏对所有人类转录异型的哈普洛型解决的功能预测.
研究的目的:
- 介绍HapScoreDB,一个全面的数据库,提供蛋白质语言模型衍生的人类蛋白质单质类型的健身分数.
- 通过将哈普洛型信息与预测的蛋白质功能相结合,促进对遗传变异的解释.
主要方法:
- 利用GENCODE和Ensembl注释与1000个基因组项目的分阶段变体数据.
- 使用最先进的蛋白质语言模型计算了超过13万种不同的蛋白质单元型的健身分数.
- 开发了一个用户友好的网络界面,用于数据探索,可视化和下载.
主要成果:
- 含有癌症基因组广泛协会研究 (GWAS) 变异的哈普洛类型显示预测适应性显著降低.
- 预测的健身分数的变化在同一转录的单个类型中突出了已知的癌症基因,表明功能重要性.
- 该数据库包括> 18,000 个基因,78,000 个转录和> 94,000 个编码变体.
结论:
- 哈普斯科尔DB提供了一种新的资源,用于理解人类遗传变异在哈普类型水平上的功能后果.
- 该平台能够实现先进的变体解释,同型优先级和人口规模的功能基因组学.
- 它将现实世界的人类遗传数据与尖端蛋白质建模连接起来,以增强生物洞察力.
相关概念视频
Protein Families
16.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.6K
Gene Families
9.8K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
9.8K
Protein-protein Interfaces
14.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.4K
Globular and Fibrous Proteins
46.7K
Many proteins can be classified into two distinct subtypes - globular or fibrous. These two types differ in their shapes and solubilities.
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
46.7K
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K


