阿斯卡里斯:基于单氨基酸变异的位置特征注释和基于蛋白质结构的表示
Fatma Cankara1,2,3, Tunca Doğan1,4,5
1Biological Data Science Laboratory, Dept. of Computer Engineering, Hacettepe University, Ankara, Turkey.
Computational and structural biotechnology journal
|October 12, 2023
概括
我们开发了ASCARIS,这是一种使用功能注释来表示单氨基酸变异 (SAV) 的新方法. 这种方法有助于预测变异效应,并为遗传疾病研究建立综合模型.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 基因组变异可能会对蛋白质功能和生物过程产生负面影响,因此需要方法来了解它们的影响.
- 对序列变异的准确分析对于开发遗传疾病治疗方法至关重要.
- 现有的计算方法需要改进,以准确地表示和数据驱动的分析序列变化.
研究的目的:
- 引入ASCARIS (单个氨基酸变异的注释和基于结构的表示),这是单个氨基酸变异 (SAV) 的定量表示方法.
- 为了能够预测SAV的功能影响和构建多omics整合模型.
- 通过功能注释,提供可重复使用的SAV数值表示.
主要方法:
- 阿斯卡里斯集成SAV位置与30个位置特征注释 (例如,活跃站点,绑定区域) 和结构/物理化学性质.
- 该方法侧重于使用功能注释创建SAV的数值表示.
- 使用ASCARIS表示方式训练了变量效应预测模型.
主要成果:
- 统计分析证实,ASCARIS使用的个别特征包含有关变异后果的信息.
- 在变异效应预测中,ASCARIS表现出与最先进的预测器相比具有竞争力和互补性.
- 废除研究和比较验证了ASCARIS表示的疗效.
结论:
- 阿斯卡里斯为代表SAV提供了一个新的,功能性的视角.
- 该方法可以独立使用或与其他SAV分析方法结合使用.
- 作为一个程序工具和一个网络服务,ASCARIS可用于更广泛的应用.
更多相关视频
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
11.0K
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
7.3K
相关概念视频
Protein Organization
6.5K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.5K
Protein and Protein Structure
79.7K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
79.7K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K
Protein and Protein Structures
10.5K
10.5K
Amino acids
89.1K
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible...
89.1K
