Related Experiment Video
Updated: Sep 11, 2025

13:51
Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
20.1K
Covariance Assisted Multivariate Penalized Additive Regression (CoMPAdRe)
Neel Desai1, Veerabhadran Baladandayuthapani2, Russell T Shinohara1
1Division of Biostatistics, University of Pennsylvania.
Summary
We introduce Covariance Assisted Multivariate Penalized Additive Regression (CoMPAdRe), a new method for selecting and estimating sparse additive models. This approach improves variable selection and estimation efficiency by accounting for inter-response correlation in multivariate data.
Area of Science:
- Statistics
- Bioinformatics
- Genomics
Background:
- Multivariate statistical modeling is crucial for analyzing complex biological data.
- Accounting for inter-response correlation can enhance model accuracy.
- Existing methods often analyze responses independently, potentially missing important relationships.
Purpose of the Study:
- To develop a novel method for simultaneous selection and estimation of multivariate sparse additive models.
- To improve variable selection accuracy and estimation efficiency by incorporating joint estimation of residual structures.
- To apply the method to analyze protein-mRNA expression levels in breast cancer pathways.
Main Methods:
- Covariance Assisted Multivariate Penalized Additive Regression (CoMPAdRe) method.
- Simultaneous selection of null, linear, and non-linear effects for each predictor.
- Joint estimation of sparse residual structure among responses.
- Computationally efficient parallel processing across responses.
Main Results:
- CoMPAdRe demonstrates improved estimation efficiency and selection accuracy compared to single-response approaches.
- Gains are more significant in settings with moderate signal relative to noise.
- Non-linear mRNA-protein associations were identified in several breast cancer pathways (Core Reactive, EMT, PIK-AKT, RTK).
Conclusions:
- Joint multivariate modeling accounting for inter-response correlation offers substantial benefits in statistical analysis.
- CoMPAdRe provides a powerful tool for characterizing complex biological associations, such as mRNA-protein relationships.
- The findings contribute to a better understanding of molecular pathways in breast cancer.
Related Concept Videos
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Correlation and Regression
1.9K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.9K
Regression Analysis
6.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.0K
Coefficient of Correlation
6.4K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
6.4K
Variation
7.2K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
7.2K
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K

