通过引导自适应收缩利用外部信息来改善高维回归设置中的变量选择
Mark A van de Wiel1, Wessel N van Wieringen1,2
1Department of Epidemiology and Data Science, Amsterdam Public Health Research Institute, Amsterdam University Medical Centers, Amsterdam, The Netherlands.
The international journal of biostatistics
|September 17, 2025
概括
导向自适应收缩方法利用外部共同数据来增强高维,低样本尺寸设置中的变量选择. 这种方法通过使用补充信息来调整收缩参数来提高预测准确性,特别是在基因组学中.
科学领域:
- 统计 统计 统计 统计
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
背景情况:
- 具有较小样本大小的高维数据带来了重大的变量选择挑战.
- 外部信息,称为"共同数据",可以提高变量选择的准确性.
- 共同数据,如变量分组或先前的p值,在基因组学中非常丰富.
研究的目的:
- 审查使用共同数据进行改进变量选择的引导自适应收缩方法.
- 讨论共同数据在预测模型中的技术方面和适用性.
- 为了比较引导收缩与其他方法,如稀疏组激光.
主要方法:
- 导向自适应收缩方法的审查.
- 使用共同数据调整收缩参数.
- 与稀疏组-拉索进行变量选择的比较.
- 整合共同数据学习者和尖端和平板优先级,以实现"自己做"的实现.
主要成果:
- 引导自适应收缩方法有效地使用共同数据来增强变量选择.
- 该方法在整合不同类型的共同数据方面表现出了多功能性.
- 通过DIY实施在遗传学研究中改进变量选择的演示.
结论:
- 引导自适应收缩提供了一个强大的框架,用于可变选择与共同数据.
- 这些方法适用于各种领域,特别是基因组学.
- 为研究人员提供了实际实施指南.
相关概念视频
Regression Toward the Mean
6.9K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.9K
Multiple Regression
3.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.8K
Regression Analysis
8.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.1K
Survival Tree
389
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
389
Variation
7.7K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
7.7K
Outliers and Influential Points
6.1K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.1K

