发现和分析蛋白质中的重复性和低复杂性架构及其保存的进化关系,使用自相对论点图
Maria W Górna1, Matthew Merski2
1Structural Biology Group, Biological and Chemical Research Centre, Faculty of Chemistry, University of Warsaw, Warsaw, Poland.
Methods in molecular biology (Clifton, N.J.)
|November 14, 2024
概括
自同学点图分析是一种独立于模型的方法,用于识别蛋白质序列的重复. 这种技术有效地检测出不同保存的重复,并识别出蛋白质家族中的保存模式.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 结构生物学 结构生物学
背景情况:
- 具有重复序列和低复杂度区域的蛋白质带来了分析挑战.
- 识别和描述这些重复的元素对于理解蛋白质的功能和进化至关重要.
研究的目的:
- 将自我同质点图分析作为识别蛋白质序列重复的强有力的方法.
- 为蛋白质重复和保存模式的统计和视觉识别提供准则.
主要方法:
- 使用自同质点图分析可视化蛋白质内的序列重复.
- 应用统计标准来识别蛋白质的重复,包括那些保存率较低的蛋白质.
- 采用等级聚类和Jaccard索引来分析蛋白质家族中的重复架构和保存.
主要成果:
- 点图有效地识别蛋白质重复的数量,长度和位置,而不需要预定义的模型.
- 该方法成功检测了低序列保存和非并列重复的重复.
- 对点图的视觉检查揭示了蛋白质家族内保存的模式,可以通过雅卡德指数量化.
结论:
- 自同学点图分析是一种多功能且独立于模型的工具,用于表征蛋白质重复.
- 这种方法有助于在蛋白质序列中发现新的重复结构和保存特征.
- 该方法有助于理解蛋白质中重复元素的进化和功能意义.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Gene Evolution - Fast or Slow?
7.0K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.0K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


