一种新的容器大小指数方法,用于从材料表征数据集的多式联运数据集的统计分析
Tao Jiang1, Shengmin Luo1,2, Dongfang Wang1
1Department of Civil and Environmental Engineering, University of Massachusetts Amherst, Amherst, MA, 01003, USA.
Scientific reports
|July 5, 2023
概括
一种新的统计数据分类方法,分类大小指数 (BSI),客观地优化了直方图分类大小. 该方法改善了用于材料表征的多模式数据集的解卷,并准确地确定了概率密度函数.
科学领域:
- 材料科学 材料科学 材料科学
- 统计分析 统计分析
- 数据挖掘 数据挖掘
背景情况:
- 立体图的构造对于材料表征中的数据分析至关重要.
- 客观地确定最佳垃圾箱大小是一个重大挑战.
- 现有的方法经常与多式联运数据集和过度装配作斗争.
研究的目的:
- 引入一种新的,客观的统计数据分类方法,称为分类大小指数 (BSI).
- 为了实现合理的直方图构造,以便有效地解卷多式联运数据集.
- 准确确定材料表征中的潜在概率密度函数.
主要方法:
- 开发了一种标准化的基于错误的统计数据组合方法 (BSI).
- 将BSI应用于合成和现实数据集 (岩石弹性,粘土悬浮).
- 与其他广泛使用的捆绑方法比较BSI表现.
主要成果:
- BSI成功地确定了直方图的最佳容器大小.
- 该方法准确地解构了多式联运数据集,并确定了概率密度函数.
- 衡量标准表现优于其他方法,产生更高的衡量标准值和更小的规范化标准误差.
- BSI有效地惩罚了过度装配,并确定了数据集中的模式数量.
结论:
- BSI方法提供了一个客观而准确的方法,用于材料表征的数据对接.
- 它增强了复杂的多式联运数据集的解卷.
- BSI为确定概率密度函数并避免过拟合提供了强大的解决方案.
相关概念视频
One-Way ANOVA: Unequal Sample Sizes
5.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.8K
Sample Size Calculation
3.6K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
3.6K
Statistical Methods to Analyze Parametric Data: ANOVA
440
Analysis of Variance, or ANOVA, is a powerful statistical technique used to analyze parametric data, primarily in research and experimental studies. It's designed to compare the means of two or more groups, assisting researchers in identifying any significant differences between these group means. There are two main types of ANOVA based on the complexity of the analysis: one-way and two-way.
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
440
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K
Statistical Analysis: Overview
6.7K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.7K
Empirical Method to Interpret Standard Deviation
5.3K
The empirical rule, also known as the three-sigma rule, allows a statistician to interpret the standard deviation in a normally distributed dataset. The rule states that 68% of the data lies within one standard deviation from the mean, 95% lies within two standard deviations from the mean, and 99.7% lies within three standard deviations from the mean. Additionally, this rule is also called the 68-95-99.7 rule.
This rule is used widely in statistics to calculate the proportion of data values...
This rule is used widely in statistics to calculate the proportion of data values...
5.3K


