相关实验视频
Updated: Jul 10, 2025

08:57
Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
15.9K
一种基于序列的进化距离方法,用于对高度分离的蛋白质进行基因分析
Wei Cao1, Lu-Yun Wu1, Xia-Yu Xia1
1Key Laboratory of Ministry of Education for Protein Science, School of Life Sciences, Tsinghua University, Beijing, 100084, China.
Scientific reports
|November 21, 2023
概括
一个新的序列距离 (SD) 算法改进了对分离的蛋白质序列的基因分析. 这种方法准确地揭示了进化关系和蛋白质结构,为现有工具提供了更快的替代方案.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 进化生物学 进化生物学
背景情况:
- 对高度分离的蛋白质序列的基因分析是具有挑战性的,因为当前方法的局限性.
- 准确的进化关系推断对于理解蛋白质家族的进化和功能至关重要.
研究的目的:
- 开发一种新的基于序列的进化距离算法,称为序列距离 (SD).
- 为了提高对不同蛋白质序列的遗传学分析的准确性和效率.
- 为探索蛋白质家族内部和之间进化关系提供可靠的工具.
主要方法:
- 开发了序列距离 (SD) 算法,结合了蛋白质序列中的站点对站点相关性.
- 应用SD来分析蛋白质超级家族中的进化关系.
- 与基于结构信息的SD衍生类遗传树进行比较.
主要成果:
- SD有效地区分了蛋白质家族内部和蛋白质家族之间的进化关系,即使与<20%的序列相同性.
- 由SD生成的家族遗传树与基于结构信息的树密切结合.
- SD在单个CPU上快速计算进化距离 (每千对秒),性能优于结构预测方法.
结论:
- 序列距离 (SD) 算法为分离的蛋白质的遗传学分析提供了显著的进步.
- SD为进化研究提供了一个更准确,更可靠,更高效的计算工具.
- 这种方法增强了对蛋白质进化和结构相似性的理解.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K

