复杂添加剂图形推断的低准确性从f统计数据中推断出来
Lauren E Frankel1,2, Cécile Ané3
1Department of Botany, University of Wisconsin-Madison, Madison, WI 53706, United States.
Genetics
|December 2, 2025
概括
F-统计学准确地推断出简单的进化网络,但与复杂的网络作斗争. 主要树结构可靠地恢复,但在解释复杂的网络细节时建议谨慎.
科学领域:
- 进化生物学是进化的生物学.
- 人口遗传学 人口遗传学
- 人类遗传学 是一个学科.
背景情况:
- F统计被广泛用于分析种群和进化系之间的杂交,混合和内进化.
- 推断复杂的进化史,特别是涉及网状物进化的进化史,仍然是植物遗传学的一个重大挑战.
研究的目的:
- 评估f-统计学在推断不同复杂度的遗传学网络结构中的准确性.
- 通过使用f统计学来确定影响网络推理可靠性的因素.
主要方法:
- 利用模拟来测试f统计的性能在一系列的族系网络复杂性.
- 评估了网络特征 (如循环大小,网格数量和分子时钟违规) 对推断准确性的影响.
主要成果:
- F-统计学准确地恢复了单个网状或大循环 (≥4个节点) 中的网状网络结构.
- 对于复杂的网络来说,推理准确性很差,特别是那些以小周期 (3节点) 的网状结构,增加的网状结构数量或更高的网络水平的网络.
- 无论网络的复杂性如何,主要的家族遗传树拓一直可靠地恢复.
- 违反分子时钟显著降低了网络推断的准确性,并增加了网络事件的错误拒绝.
结论:
- 网络复杂性,特别是小周期的存在,是限制使用f统计学准确推断网状细胞进化的关键因素.
- 可识别性问题可能是简单与复杂网络的可恢复性差异的基础.
- 虽然主要树组件可靠地估计,但推断详细的网络结构,特别是复杂的网络结构,需要仔细考虑和验证.
- 建议包括评估多个得分最高的网络,并评估研究系统内的速率变化.
相关概念视频
F Distribution
8.8K
The F distribution was named after Sir Ronald Fisher, an English statistician. The F statistic is a ratio (a fraction) with two sets of degrees of freedom; one for the numerator and one for the denominator. The F distribution is derived from the Student's t distribution. The values of the F distribution are squares of the corresponding values of the t distribution. One-Way ANOVA expands the t test for comparing more than two groups. The scope of that derivation is beyond the level of this...
8.8K
Identifying Statistically Significant Differences: The F-Test
3.0K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
3.0K
Accuracy and Errors in Hypothesis Testing
550
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
550
Expected Frequencies in Goodness-of-Fit Tests
7.1K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
7.1K
Hypothesis Test for Test of Independence
7.4K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
7.4K
Goodness-of-Fit Test
8.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
8.1K


