评估序列相似性网络的客观标准有助于划分蛋白质家族的序列空间
Bastian Volker Helmut Hornung1,2, Nicolas Terrapon1,2
1Aix Marseille Université, CNRS, UMR 7257 AFMB, Marseille, France.
PLoS computational biology
|August 16, 2023
概括
计算式蛋白质注释面临着大量基因组数据的挑战. 本研究介绍了序列相似性网络 (SSN) 中的近距离中心性,作为定义蛋白质子家族的数据驱动方法,改善功能注释.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 基因组数据洪水需要先进的计算蛋白质注释方法.
- 蛋白质家族需要分为子家族进行细分,以便精确的功能注释和发现.
- 手动定义子家族是耗时的,对于大数据集来说是不可扩展的.
研究的目的:
- 开发可扩展和可复制的蛋白质子家族定义方法.
- 确定一个数据驱动的标准来评估序列相似性网络 (SSN) 子家族的分配.
- 增强功能注释,指导发现新型蛋白质功能.
主要方法:
- 利用序列相似性网络 (SSN) 来分析大型蛋白质家族.
- 研究了网络属性,重点关注接近的中心性,以评估子家族的分配.
- 应用了接近中心性来定义四个碳水化合物活性酶 (CAZy) 家族中的子家族.
主要成果:
- 亲密的中心性与专家策展人的决策在子家族任务中的高度相似.
- 这个网络属性可以表示多个子家族级别或没有明确的子家族.
- 使用接近中心性的子家族创建在CAZy家族中提供了更细致的功能注释.
结论:
- 接近中心性为蛋白质子家族定义提供了一个数据驱动的,可扩展的方法.
- 这种方法有助于识别未具特征的子家族,以便在未来进行研究.
- 该方法提高了大型蛋白质家族中的功能注释的精度.
更多相关视频
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
7.3K
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
2.7K
相关概念视频
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Gene Evolution - Fast or Slow?
7.2K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.2K
Gene Families
8.9K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.9K
