超级分区:快速,灵活和可解释的大规模数据减少RR
Katelyn J Queen1, Malcolm Barrett2, Joshua Millstein1
1Department of Population and Public Health Sciences, University of Southern California, Los Angeles, California, United States.
PeerJ
|January 31, 2025
概括
超级分区为大型,复杂的数据集提供了可扩展的数据减少方法. 这种对 Partition 算法的近似增强了用于高维数据分析的计算可操作性.
科学领域:
- 数据科学数据科学数据科学
- 计算统计学 计算统计学
背景情况:
- 随着数据大小和复杂性的增加,需要先进的数据缩小技术.
- 现有的方法需要灵活的信息保存量化.
研究的目的:
- 介绍超级分区,这是分区算法的可扩展的近似.
- 允许灵活规范每个特征的最小信息捕获.
主要方法:
- 使用Genie,一个快速的等级聚类算法,用于初始的超级分区形成.
- 将分区算法应用于生成的子集,以实现计算可处理性.
主要成果:
- 展示高维数据集的数十万个特征的可扩展性.
- 实现合理的计算时间,以大规模减少数据.
结论:
- 超级分区提供了一种高效和灵活的数据缩小方法.
- 该方法适用于有效处理复杂和高维数据集.
相关概念视频
Introduction to R
226
R is a powerful software environment for statistical computing and graphics. Originating as an implementation of the S language, developed at Bell Laboratories, R has evolved into a robust, open-source statistical software favored by statisticians and data scientists worldwide. Its comprehensive suite includes data manipulation, calculation, and graphical display capabilities, making it versatile for data analysis and visualization. Its programming language is at the core of R's...
226
One-Way ANOVA: Equal Sample Sizes
3.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.2K
One-Way ANOVA: Unequal Sample Sizes
5.7K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.7K
Extraction: Partition and Distribution Coefficients
1.7K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
1.7K
Sampling Plans
165
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
165
Friedman Two-way Analysis of Variance by Ranks
137
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
137


