除了中心倾向:如果我们同意离散的植被群落不存在,我们应该研究其他集群方法吗?
Mark G Tozer1,2, David A Keith2
1NSW Department of Planning and Environment Parramatta New South Wales Australia.
Ecology and evolution
|November 29, 2023
概括
植物分类稳定性通过考虑样本相互连接性和距离中心点的距离而得到改善. 删除弱连接样本可以提高分类的稳定性,特别是较少的集群.
科学领域:
- 生态生态学 生态生态学
- 数据科学数据科学数据科学
- 计算生物学 计算生物学
背景情况:
- 强大的植被分类对于理解生态模式至关重要.
- 当前的集群方法通常会产生对数据和算法敏感的不稳定分类解决方案.
- 不稳定性传统上归因于噪音,定义为偏离集群的中心趋势.
研究的目的:
- 调查植被分类中不稳定的原因.
- 为了比较中央倾向与相互连接模型在分类稳定性的预测能力.
- 评估采样强度和异常值去除对分类稳定性的影响.
主要方法:
- 在5次代中模拟了采样强度的增量增加.
- 使用基于中心趋势和图形理论互连性模型的算法评估分类稳定性.
- 采用后勤回归来模型样本组的变化,基于距离心心和相互连接程度的距离.
主要成果:
- 样本相互连接的程度是分类不稳定的更强有力的预测因素,而不是距离中心点的距离.
- 删除弱相互连接的样本导致了更稳定的植被分类.
- 随着集群数量的增加,分类稳定性下降,异常值去除的好处减少.
结论:
- 来自连续植被数据的集群本质上是不稳定的.
- 专注于相互连接的图形理论方法为连续数据提供了更稳定的分类.
- 从大型区域数据集中寻求稳定,精细的分类可能是不现实的.
相关概念视频
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Sampling Plans
187
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
187
Distribution and Dispersion
21.8K
To understand intra-specific interactions in populations, scientists measure the spatial arrangement of species individuals. This geographic arrangement is known as the species distribution or dispersion. Highly territorial species exhibit a uniform distribution pattern, in which individuals are spaced at relatively equal distances from one another. Species that are highly tied to particular resources, such as food or shelter, tend to concentrate around those resources, and thus exhibit a...
21.8K
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Survival Tree
87
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
87
What is Central Tendency?
14.8K
Descriptive statistics describe or summarize relevant characteristics of a sample and aid in the analysis of data of interest. When analyzing large quantities of data and developing an inference, one needs to identify a value representative of the entire data set. Characteristics such as central tendency, extreme values, range of measurements, or the most repeated value can help better understand the data.
The central tendency is the most conventionally used data characteristic. It is a...
The central tendency is the most conventionally used data characteristic. It is a...
14.8K


