相关实验视频
Updated: Jul 19, 2025

12:00
A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
35.4K
数据特定的替代模型改善了基于蛋白质的遗传学
João M Brazão1, Peter G Foster2, Cymon J Cox1
1Centro de Ciências do Mar, Universidade do Algarve, Faro, Algarve, Portugal.
PeerJ
|August 14, 2023
概括
对蛋白质进行准确,数据特定的氨基酸替代模型的估计对于植物遗传学分析至关重要. 这项研究发现,数据特定模型,特别是来自IQ-TREE和P4 (最大概率) 的模型,比实证模型提供更高的准确性和效率.
科学领域:
- 计算生物学 计算生物学
- 人类遗传学 是一个学科.
- 生物信息学是一种生物信息学.
背景情况:
- 估计准确的氨基酸替代模型对于遗传学分析至关重要,但在计算上具有挑战性.
- 现有的经验模型可能不能充分代表特定的蛋白质数据集,可能导致不准确的进化推断.
研究的目的:
- 评估五种用于估计特定数据的氨基酸替代模型的计算效率和准确性.
- 将特定数据模型的性能与常用的实证模型 (cpREV,WAG) 进行比较.
- 证明使用特定数据模型在已公布数据集的基因分析中的好处.
主要方法:
- 使用已知的模型和树来生成模拟的蛋白质对齐.
- 使用五种方法 (Codeml,FastMG,IQ-TREE,P4-ML,P4-贝叶斯) 来估计特定数据模型.
- 在数据特定和实证模型与模拟模型之间进行了最大概率得分和树拓的比较.
主要成果:
- 数据特定模型显著优于实证模型,显示更好的匹配和推断更准确的树.
- IQ-TREE和P4 (最大概率) 方法产生了与模拟模型在统计学上无法区分的树木.
- 用IQ-TREE估计模型重新分析已发布的数据集,发现超过一半的案例数据匹配度有所改善,推断树长,拓结构改变.
结论:
- 数据特定的氨基酸替代模型提供了比实证模型更准确的蛋白质进化表现.
- 建议使用IQ-TREE和P4 (最大概率),因为它们在估计这些模型中的准确性和效率.
- 计算负担和软件的可用性不再是为生成先进的数据特定模型进行遗传学分析的重大限制.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Phylogenetic Trees
45.5K
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.
45.5K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Gene Evolution - Fast or Slow?
7.2K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.2K
Phylogeny
44.4K
Phylogeny is concerned with the evolutionary diversification of organisms or groups of organisms. A group of organisms with a name is called a taxon (singular). Taxa (plural) can span different levels of the evolutionary hierarchy. For instance, the group containing all birds is a taxon (comprising the class Aves), and the group of all species of daisies (the genus Bellis) is a taxon. Phylogenies can likewise include just one genus (i.e., depict species relationships) or span an entire kingdom.
44.4K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K

