TransEFVP:基于蛋白质序列嵌入融合的人类致病变体预测的两阶段方法
Zihao Yan1, Fang Ge2, Yan Liu3
1School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, PR China.
Journal of chemical information and modeling
|February 9, 2024
概括
我们开发了TransEFVP,这是一个使用蛋白质语言模型和神经网络的计算工具,用于预测与疾病相关的单氨基酸变异 (SAV). 这种方法可以准确地识别有害的遗传变化,进步精准医学.
科学领域:
- 基因组学就是基因组学.
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
背景情况:
- 了解单氨基酸变异 (SAV) 对分子生物学,进化和疾病至关重要.
- 识别有害变异是精准医学的一个关键挑战.
研究的目的:
- 引入TransEFVP,一种用于预测与疾病相关的SAVs的新型计算方法.
- 为了利用大规模的蛋白质语言模型嵌入和变换器神经网络用于病原性预测.
主要方法:
- TransEFVP采用两级架构,将变压器编码器结合起来进行特征融合.
- 在减小维度后,应用支持向量机 (SVM) 模型来量化病原性.
主要成果:
- 在盲测试数据上,TransEFVP实现了0.751的马修斯相关系数,0.846的F1得分和0.871的AUC.
- 绩效指标超过了现有的最先进的方法.
结论:
- 作为一种SAV病原性预测工具,TransEFVP表现出高准确性和有效性.
- 该方法显示出在精密医学和了解疾病机制方面的应用潜力.
相关概念视频
Tagging and Fusion Proteins
6.6K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.6K
Cotranslational Protein Translocation
7.4K
Translocation of proteins across membranes is an ancient process that occurs even in bacteria and archaebacteria. In fact, the components of the translocation machinery are still conserved between prokaryotes and eukaryotes.
Sec61 channel partners for cotranslational translocation
During cotranslational translocation, the Sec61 channel partners with the signal recognition particle (SRP), the signal recognition particle receptor (SR), and the ribosomes to transport the nascent polypeptide chain...
Sec61 channel partners for cotranslational translocation
During cotranslational translocation, the Sec61 channel partners with the signal recognition particle (SRP), the signal recognition particle receptor (SR), and the ribosomes to transport the nascent polypeptide chain...
7.4K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Insertion of Multi-pass Transmembrane Proteins in the RER
8.0K
The rough ER membrane synthesizes, assembles, and embeds transmembrane proteins in diverse topologies. These proteins function as transporters or channels and can remain in the ER membrane or are sent to the Golgi complex, lysosome, and cell membrane.
The multipass transmembrane proteins are the type IV integral membrane proteins with multiple topogenic sequences determining their spatial arrangement in the ER membrane. Nearly all multipass proteins lack a cleavable signal sequence and use...
The multipass transmembrane proteins are the type IV integral membrane proteins with multiple topogenic sequences determining their spatial arrangement in the ER membrane. Nearly all multipass proteins lack a cleavable signal sequence and use...
8.0K
Single Nucleotide Polymorphisms-SNPs
15.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K
Multi-pass Transmembrane Proteins and β-barrels
5.3K
In multi-pass transmembrane proteins, the polypeptide chain crosses the membrane more than once. The transmembrane polypeptide chain either forms an α-helix or β-strand structure. α-Helix containing multi-pass transmembrane proteins are ubiquitous, whereas β-strand containing ones are mainly found in gram-negative bacteria, mitochondria, and chloroplasts.
α-Helix containing multi-pass transmembrane proteins
Multi-pass transmembrane proteins such as...
α-Helix containing multi-pass transmembrane proteins
Multi-pass transmembrane proteins such as...
5.3K


