EMMVEP:多源特徴量融合に基づくタンパク質ミスセンスバリアント効果予測のためのアンサンブル手法
Huiling Zhang1, Junwen Huang2, Yuetong Li1
1College of Mathematics and Information, College of Software Engineering, South China Agricultural University, Guangzhou, 510642, China.
Interdisciplinary sciences, computational life sciences
|February 27, 2026
まとめ
EMMVEPは、タンパク質のミスセンス変異の効果を予測する新しい計算ツールです。この手法は、病原性バリアントと良性バリアントを正確に区別し、研究および臨床設定における遺伝子バリアントの解釈を支援します。
科学分野:
- ゲノミクス; 計算生物学; タンパク質科学
背景:
- ミスセンス変異は、タンパク質の機能を変化させる可能性のある一般的な遺伝的バリエーションであり、病原性バリアントと良性バリアントを区別する上で課題となっています。ミスセンス変異効果の正確な予測は、遺伝性疾患の理解と臨床的意思決定の指針にとって重要です。
研究 の 目的:
- タンパク質のミスセンス変異の機能的影響を予測するためのアンサンブルベースの計算アプローチであるEMMVEPを導入すること。既存のバリアント効果予測手法と比較してEMMVEPのパフォーマンスを評価すること。
主な方法:
- EMMVEPは、タンパク質配列情報、AlphaFoldからの物理化学的特性、gnomADからのアレル頻度を含む多様な特徴量を統合します。変異病原性予測のためのアンサンブルモデルを構築するために、カテゴリカルブースティングが採用されています。
主要な成果:
- EMMVEPは、ベンチマークデータセットで高いパフォーマンスを達成し、曲線下面積(AUC)は0.907、適合率-再現率曲線下面積(AUPR)は0.879でした。この手法は、20の一般的なバリアント効果予測ツールを上回りました。19,233のヒト遺伝子にわたる2億1600万以上の潜在的なアミノ酸置換に対する病原性確率が提供されます。
結論:
- EMMVEPは、ミスセンス変異効果を予測するための堅牢で正確な方法を提供し、遺伝子バリアントの解釈を向上させます。このツールは、病原性変異を特定するための貴重な洞察を提供し、研究と臨床応用の両方に大きな影響を与えます。
関連する概念動画
Conservation of Protein Domains Over Different Proteins
14.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.8K
Tagging and Fusion Proteins
8.6K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
8.6K
Multi-species Conserved Sequences
4.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.9K
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
Improving Translational Accuracy
3.7K
3.7K
Comparing Copy Number Variations and SNPs
18.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.9K


