对算法的中立比较,以尽量减少对高维变量选择的L0处罚
1Institute of Medical Statistics, Center for Medical Data Science, Medical University of Vienna, Vienna, Austria.
Biometrical journal. Biometrische Zeitschrift
|July 8, 2023
概括
新的算法改善了在高维数据中稀疏的模型选择的L0惩罚最小化. 模拟和现实世界遗传数据分析显示,这些先进的变量选择技术的性能和效率有所提高.
科学领域:
- 统计 统计 统计 统计
- 计算生物学 计算生物学
- 遗传学 遗传学 是一个
背景情况:
- 对于在高维数据中稀疏的模型选择,L0惩罚方法提供了理论上的优势.
- 现有的方法,如修改的贝叶斯信息标准 (mBIC, mBIC2) 控制错误率,但由于NP-hard L0最小化而面临计算挑战.
- 由于计算容易,像LASSO这样的凸变替代品很受欢迎,但L0方法正在看到算法进步.
研究的目的:
- 为了比较最近设计的算法的性能,以最大限度地减少基于L0的选择标准.
- 评估不同L0最小化算法的统计属性和计算运行时间.
- 为了证明这些算法的实际应用在表达式定量特征位置 (eQTL) 映射中.
主要方法:
- 在各种场景中进行模拟研究,灵感来自遗传关联研究.
- 通过各种L0最小化算法获得的选择标准值的比较.
- 分析所选模型和算法运行时间的统计特征.
主要成果:
- 模拟结果提供了对L0最小化算法性能的全面比较.
- 评估算法的统计性质和计算效率.
- 该研究说明了这些算法在真实eQTL映射数据集上的实际实用性.
结论:
- 最近的算法发展使得L0惩罚最小化在计算上更加可行.
- 该研究为选择适当的算法提供了有价值的见解,用于在高维设置中稀疏的模型选择.
- 这些先进的方法显示出在遗传关联研究和eQTL映射中的应用的前景.
相关概念视频
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
81
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
81
Quantifying and Rejecting Outliers: The Grubbs Test
1.7K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.7K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Comparing the Survival Analysis of Two or More Groups
227
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
227
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Calibration Curves: Linear Least Squares
1.4K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
1.4K


