一种监督除草方法,以集群高维预测器,并应用于就业市场分析
Yuyang Li1, Jianxin Bi2, Jingyuan Liu2
1Department of Statistics, Iowa State University, Ames, IA, USA.
Journal of applied statistics
|December 4, 2024
概括
本研究介绍了一种监督除草算法,用于聚类高维数据. 该方法有效地识别稀疏集群并预测响应变量,优于现有方法.
科学领域:
- 数据挖掘是一种数据挖掘.
- 生物信息学是一种生物信息学.
- 文本分析 文本分析
背景情况:
- 在诸如文本挖掘和生物数据分析等领域,聚类高维预测器至关重要.
- 标准集群方法可以识别预测器中的模式,但不能预测响应变量.
- 现有的方法往往侧重于聚类或预测,而不是两者兼而有之.
研究的目的:
- 开发一个监督除草算法,同时检测稀疏集群和捕获预测效应.
- 在高维数据分析中解决聚类和预测的双重要求.
- 引入一种用于监督预测因素选择和聚类的新方法.
主要方法:
- 一个代的特征选和连贯性评估程序.
- 反向消除不重要的预测因素以形成嵌套集.
- 蒙特卡洛模拟用于评估有限样本的性能.
- 应用到职位描述数据集来分析关键词对工资的影响.
主要成果:
- 拟议的监督除草算法展示了与现有的专业方法相比较的聚类和预测性能.
- 该方法有效地识别了影响响应变量的重要预测组.
- 模拟结果证实了算法的有限样本性能.
结论:
- 监督除草算法为聚类和预测在高维数据中的双重问题提供了强大的解决方案.
- 这种方法通过结合预测能力来增强聚类的实用性.
- 该方法为复杂数据集中的特征重要性和关系提供了有价值的见解.
相关概念视频
Survival Tree
60
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
60
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K


