Related Experiment Video
Updated: Mar 15, 2026

04:35
Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
3.8K
Effect of correlation on covariate selection in linear and nonlinear mixed effect models
1Astellas Pharma Global Development, 1 Astellas Way, N2.292, Northbrook, IL, 60062, USA.
Pharmaceutical Statistics
|September 2, 2016
Summary
Correlated covariates can falsely appear as significant predictors in population pharmacokinetic analyses. This can lead to the selection of incorrect covariates, impacting drug development and analysis.
Area of Science:
- Pharmacometrics
- Population Pharmacokinetics
- Statistical Modeling
Background:
- Covariate selection is crucial in population pharmacokinetic (PopPK) analyses for understanding drug behavior.
- High correlation between covariates can complicate accurate selection in mixed-effects models.
- Previous studies have not fully elucidated the impact of correlated covariates on model selection.
Purpose of the Study:
- To investigate the effect of correlated covariates on covariate selection in linear and nonlinear mixed-effects models.
- To determine how correlated covariates can be mistakenly identified as significant predictors.
- To explain variability in covariate identification across different PopPK analyses.
Main Methods:
- Monte Carlo simulations were used to generate concentration-time profiles with a single true covariate affecting oral clearance (CL/F).
- Demographic data from the National Health and Nutrition Examination Survey III database were utilized.
- Population pharmacokinetic models were built using univariate covariate analysis, comparing models with and without covariates, and using likelihood ratio tests (LRT) or AIC for selection.
Main Results:
- Highly correlated covariates (e.g., weight and body surface area, r=0.98) led to the selection of the incorrect covariate (body surface area) in up to 20% of simulations.
- In a second simulation involving drug metabolites and QTc interval prolongation, a metabolite was incorrectly selected as a predictor when only parent drug concentrations influenced the outcome.
- These findings demonstrate that strong correlation alone can lead to the selection of a non-causal covariate.
Conclusions:
- Correlated covariates can be erroneously selected as significant predictors in population pharmacokinetic and pharmacodynamic (PopPK/PD) analyses due to statistical artifacts.
- This phenomenon can explain discrepancies in covariate identification for the same drug across different studies.
- Careful consideration of covariate correlations is essential for robust and accurate PopPK/PD model building.
Related Concept Videos
Correlation and Regression
3.9K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
3.9K
Correlation
15.6K
In statistics, two variables are said to be correlated if the values of one variable are associated with the other variable. Depending on the relationship between two variables, correlation can be of three types– positive correlation, negative correlation, and zero correlation.
Two variables, for example, a and b, are said to be positively correlated if both variables move in the same direction. In other words, a positive correlation exists between two variables, a and b, if:
Two variables, for example, a and b, are said to be positively correlated if both variables move in the same direction. In other words, a positive correlation exists between two variables, a and b, if:
15.6K
Calculating and Interpreting the Linear Correlation Coefficient
8.4K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
8.4K
Coefficient of Correlation
9.0K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
9.0K
Calibration Curves: Correlation Coefficient
5.3K
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the...
5.3K
Multiple Regression
4.3K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.3K

