HIPSTR:在TreeAnnotator X中最高独立的后部子树重建
Guy Baele1, Luiz M Carvalho2, Marius Brusselmans1
1Department of Microbiology, Immunology and Transplantation, Rega Institute, KU Leuven, Leuven, Belgium.
bioRxiv : the preprint server for biology
|December 23, 2024
概括
一种新的方法,最高独立的后部子树重建 (HIPSTR),在遗传学和遗传动力学研究中始终识别出比最大基层可信度 (MCC) 树更受高度支持的基层. 此外,HIPSTR 在TreeAnnotator X.中提供了更高的计算效率.
科学领域:
- 贝叶斯的家族遗传学
- 植物动力学是关于植物动力学的.
- 计算生物学是一种计算生物学.
背景情况:
- 总结后遗传树的后部分布在贝叶斯式遗传学和遗传动力学研究中至关重要.
- 最大分类可信度 (MCC) 树是这个总结常用的方法.
- 然而,MCC树可能并不总是能够捕捉到后部分布中的叶片的全部支.
研究的目的:
- 引入和评估一种新的共识树方法,即最高独立的后部子树重建 (HIPSTR).
- 为了比较HIPSTR与MCC树在识别高度支持的分类中的性能.
- 为 TreeAnnotator X. 中的共识树估计提供更新,更快的计算程序.
主要方法:
- 开发并实施HIPSTR算法用于共识树重建.
- 使用开源软件TreeAnnotator X的更新版本进行计算例程.
- 应用HIPSTR和MCC方法来从埃博拉病毒和SARS-CoV-2数据集中重建共识树.
主要成果:
- 在所有测试的数据集中,HIPSTR 始终产生了与 MCC 树相比具有更高支持基层的共识树.
- MCC树常常错过的树干非常高 (≥0.95) 和中等至高 (≥0.50) 后面的概率.
- HIPSTR在保留这些高度支持的叶片方面表现近乎完美,表现优于MCC.
- 与MCC相比,HIPSTR在TreeAnnotator X中也显示了有利的计算性能.
结论:
- HIPSTR是一种优于MCC的方法,用于重建时间校准的共识族系,特别是用于识别得到良好支持的分类.
- 更新后的TreeAnnotator X带有增强的计算程序,有助于有效地生成共识树.
- 需要进一步的研究来探索HIPSTR与其他新兴算法 (如CCD0-MAP) 相比的性能.
相关概念视频
Survival Tree
58
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
58
Phylogenetic Trees
45.2K
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.
45.2K
Phylogeny
43.6K
Phylogeny is concerned with the evolutionary diversification of organisms or groups of organisms. A group of organisms with a name is called a taxon (singular). Taxa (plural) can span different levels of the evolutionary hierarchy. For instance, the group containing all birds is a taxon (comprising the class Aves), and the group of all species of daisies (the genus Bellis) is a taxon. Phylogenies can likewise include just one genus (i.e., depict species relationships) or span an entire kingdom.
43.6K
Truncation in Survival Analysis
156
Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
156
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Residual Plots
4.5K
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
4.5K


