列子集选择的统计视图
1Department of Statistics, Stanford University, Sequoia Hall, 390 Jane Stanford Way, Stanford, CA 94305, USA.
概括
这项研究统一了列子集选择 (CSS) 和主要变量识别,通过最大概率估计证明了它们的等价性. 它为高维度的一致CSS建立了条件,并为其应用提供了有效的方法.
科学领域:
- 统计 统计 统计 统计
- 计算机科学 计算机科学
- 数据分析 数据分析
背景情况:
- 对于大型数据集来说,减小维度至关重要.
- 列子集选择 (CSS) 和主要变量识别是常见的方法.
- 这些方法传统上被看作是分开的.
研究的目的:
- 为了证明CSS和主要变量识别之间的等价性.
- 在一个统一的半参数最大概率模型中正式化这两种方法.
- 开发用于变量选择的高效和可靠的方法.
主要方法:
- 在半参数模型中最大概率估计.
- 在比例非对称模式下对高维数据的一致性的分析.
- 开发利用总结统计和处理缺失/被审查数据的方法.
主要成果:
- 列子集选择 (CSS) 和主要变量识别被证明是相当的.
- 建立了高维度一致CSS的条件.
- 为CSS提出了高效的算法,包括对不完整数据集的算法.
结论:
- 一个统一的理论框架将计算机科学和统计方法与变量选择联系起来.
- 提出的方法为减小维度提供了有效和一致的解决方案.
- 这些发现促进了变量选择在各种数据场景中的实际应用.
相关概念视频
Column Efficiency: Rate Theory
541
The rate theory of chromatography provides quantitative insight into the shapes and widths of elution bands. These bands are based on the random-walk mechanism governing molecular migration within a column. The Gaussian profile of chromatographic bands arises from the cumulative effect of random molecular motions as they progress through the column.
During elution, a solute molecule experiences numerous transitions between stationary and mobile phases, exhibiting irregular residence times in...
During elution, a solute molecule experiences numerous transitions between stationary and mobile phases, exhibiting irregular residence times in...
541
Column Efficiency: Plate Theory
918
Band broadening in a chromatography column is measured by its efficiency. This is determined by the number of theoretical plates (N). Theoretical plate theory states that a separation column consists of a continuous series of imaginary plates where solute equilibration occurs between stationary and mobile phases.
A higher number of theoretical plates signifies better column efficiency and improved separation capabilities. Plate height affects bandwidth and separation quality; it is inversely...
A higher number of theoretical plates signifies better column efficiency and improved separation capabilities. Plate height affects bandwidth and separation quality; it is inversely...
918
Cluster Sampling Method
12.8K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.8K
Statistical Analysis: Overview
7.4K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
7.4K
One-Way ANOVA: Equal Sample Sizes
3.5K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.5K
Contingency Table
2.6K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.6K


