根据共变量进行调整,并使用MUVR2评估机器学习中的建模适应性
Yingxiao Yan1, Tessa Schillemans2, Viktor Skantze3
1Department of Life Sciences, Chalmers University of Technology, Gothenburg, Sweden.
Bioinformatics advances
|April 22, 2024
概括
新的MUVR2框架通过改进模型评估和共变量调整来增强Omics研究中的机器学习 (ML). 该工具为生物数据分析中的预测和变量选择提供了最先进的性能.
科学领域:
- 奥米克斯研究的研究.
- 生物信息学是一种生物信息学.
- 计算生物学是一种计算生物学.
背景情况:
- 机器学习 (ML) 是Omics研究的组成部分,用于分析分子数据并确定与暴露和健康的关联.
- 现有的ML方法往往缺乏强大的框架来评估模型适应性和调整共变量,阻碍了生物解释.
- 之前的MUVR算法在预测和变量选择方面展示了最先进的性能.
研究的目的:
- 引入MUVR2框架,这是MUVR算法的一个进步.
- 通过纳入模型适应性评估和共变量调整能力来解决现有的ML方法的局限性.
- 增强ML在Omics研究中的实用性,以获得更深入的生物学见解.
主要方法:
- MUVR2集成了弹性网调整回归框架,以及部分最小方形和随机森林模型.
- 该框架采用先进的交叉验证策略,以确保最先进的性能,并最大限度地减少过度装配.
- 在MUVR2.2的弹性网建模组件中启用了共变量调整.
主要成果:
- 在各种交叉验证策略中,MUVR2在预测和变量选择方面始终实现了最先进的性能.
- 该框架有效地减少了过拟合,从而使模型评估更可靠.
- MUVR2成功地展示了使用弹性网进行协变量调整的能力,在这种情况下,部分最小正方形或随机森林没有这种功能.
结论:
- MUVR2为OMIC研究中的ML应用提供了显著的进步,提供了改进的模型评估和协变量调整.
- 包括算法,数据和教程在内的MUVR2的开源可用性,促进了科学界更广泛的采用和可重复性.
- 这一框架通过使ML模型能够考虑混因素,促进了更强大的生物解释.
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Calibration Curves: Linear Least Squares
1.3K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
1.3K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K


