强大的多结果回归与相关的共变区块使用合的LAD-lasso
Jyrki Möttönen1, Tero Lähderanta2, Janne Salonen3
1Department of Mathematics and Statistics, University of Helsinki, Helsinki, Finland.
Journal of applied statistics
|March 31, 2025
概括
这项研究引入了一种强大的融合LAD-lasso方法,用于多个结果,增强高维回归中的变量选择和估计. 该方法有效地处理非正常数据和异常值,提高模型的准确性.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 计量经济学 计量经济学
背景情况:
- 高维回归模型经常面临非正常结果分布和边缘观测的挑战.
- 同时估计和选择变量对于理解复杂数据集至关重要.
- 共同变量数据经常表现出自然的相关性结构,例如顺序或空间测量.
研究的目的:
- 提出一个可靠的融合LAD-lasso方法,设计用于多种结果.
- 在处理非正常数据和异常值时,解决现有方法的局限性.
- 为了在回归中处理相关的共同变量块,要纳入一个组聚变惩罚.
主要方法:
- 一个强大的融合最小绝对偏差-拉索 (LAD-拉索) 方法,用于多个结果.
- 实施集团融合处罚,以管理相关的共变量块,并鼓励系数相似性.
- 使用贝叶斯信息标准 (BIC) 类型的标准来选择模型.
主要成果:
- 提出的方法证明了对非正常结果分布和边缘观测的稳定性.
- 集团融合惩罚有效地处理相关的共变量结构,特别是在顺序数据中.
- 广泛的模拟证实了开发方法的特性和性能.
结论:
- 强大的融合LAD-Lasso方法为高维,多结果回归中的变量选择和估计提供了强大的工具.
- 包含一个组合惩罚的方法增强了该方法对具有固有的共变量相关性数据的适用性.
- 这种方法对现实世界的应用非常有希望,包括对偏斜的纵向数据的分析.
相关概念视频
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Correlation and Regression
1.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.2K
Regression Analysis
5.5K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.5K
Friedman Two-way Analysis of Variance by Ranks
112
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
112
Comparing the Survival Analysis of Two or More Groups
96
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
96
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


