解释变量之间的未建模相互作用的混的影响,当使用隐性变量回归模型进行推断时
Olav M Kvalheim1, Warren S Vidar2, Tim U H Baumeister3
1Department of Chemistry, University of Bergen, Norway.
概括
部分最小平方 (PLS) 回归可以产生预测模型,但由于化合物相互作用,可能会误解变量的重要性. 考虑到这些相互作用,可以改善对自然产品研究的模型解释.
科学领域:
- 化学测量 化学测量 化学测量
- 自然产品研究自然产品研究
- 生物活性的查
背景情况:
- 部分最小平方 (PLS) 回归被广泛用于分析解释变量之间的线性依赖关系的数据集.
- 在PLS模型中,变量重要性排名通常用于指导实验决策,例如识别天然产品提取物中的生物活性化合物.
- 一个常见的挑战是,化合物的数量往往超过样本的数量,导致相关度和潜在的协同或对抗相互作用.
研究的目的:
- 在处理相关变量和相互作用时,研究标准部分最小平方 (PLS) 回归在解释变量重要性方面的局限性.
- 为了证明未建模的相互作用如何在天然产品研究中导致错误的结论.
- 为准确的解释和推断提供一个改进的建模方法.
主要方法:
- 应用部分最小平方 (PLS) 回归来分析具有线性依赖性的数据集.
- 在PLS模型中包含解释变量之间的相互作用项.
- 利用选择性比率图为增强模型可视化和解释.
主要成果:
- 标准PLS模型可以产生误导性的变量重要性排名,当相互作用存在且未建模时.
- 来自自然产品研究的一个实用例子说明了由于未建模的相互作用而导致的混的后果.
- 将相互作用纳入PLS模型并使用选择性比图允许更可靠的解释.
结论:
- 在复杂的天然产品混合物中,标准PLS变量重要性可能不可靠.
- 对相互作用的计算对于在这种情况下准确解释PLS模型至关重要.
- 选择性比率图提供了一个有价值的工具,用于可视化和推断来自交互意识的PLS模型的结果.
相关概念视频
Confounding in Epidemiological Studies
151
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
151
Strategies for Assessing and Addressing Confounding
87
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
87
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Mechanistic Models: Compartment Models in Individual and Population Analysis
33
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
33
Cause and Effect
10.9K
While variables are sometimes correlated because one does cause the other, it could also be that some other factor, a confounding variable, is actually causing the systematic movement in our variables of interest. For instance, as sales in ice cream increase, so does the overall rate of crime. Is it possible that indulging in your favorite flavor of ice cream could send you on a crime spree? Or, after committing crime do you think you might decide to treat yourself to a cone?
10.9K
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K


