Related Experiment Videos
Comparison of Penalty Functions for Sparse Canonical Correlation Analysis
Prabhakar Chalise1, Brooke L Fridley
1Department of Health Sciences Research, Mayo Clinic, 200 First Street SW, Rochester, MN 55905.
Summary
Sparse canonical correlation analysis (SCCA) with the SCAD penalty and Bayesian Information Criterion (BIC) filtering is recommended for genomic data analysis, especially when variables exceed subjects.
Area of Science:
- Genomics
- Bioinformatics
- Statistical Genetics
Background:
- Canonical correlation analysis (CCA) is standard for assessing associations between variable sets.
- Traditional CCA is unsuitable for high-dimensional genomic data where variables outnumber subjects.
- High variable correlation destabilizes covariance matrices in traditional CCA.
Purpose of the Study:
- To compare four penalty functions (Lasso, Elastic-net, SCAD, Hard-threshold) for sparse canonical correlation analysis (SCCA).
- To evaluate the effectiveness of the Bayesian Information Criterion (BIC) filtering step in SCCA.
- To determine the optimal SCCA method for analyzing genomic data.
Main Methods:
- Implemented and compared Lasso, Elastic-net, SCAD, and Hard-threshold penalties for SCCA.
- Applied SCCA with and without the Bayesian Information Criterion (BIC) filtering step.
- Utilized both simulated and real genotypic and mRNA expression data for validation.
Main Results:
- SCAD penalty demonstrated superior performance in SCCA for genomic datasets.
- The addition of the BIC filtering step further enhanced the selection of relevant features.
- The combination of SCAD penalty and BIC filtering provided robust and interpretable results.
Conclusions:
- The SCAD penalty combined with BIC filtering is the preferred method for sparse canonical correlation analysis in genomics.
- This approach effectively addresses challenges posed by high-dimensional and correlated variables in genomic studies.
- The findings offer valuable guidance for researchers analyzing complex genomic datasets.
Related Concept Videos
Coefficient of Correlation
8.0K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
8.0K
Calculating and Interpreting the Linear Correlation Coefficient
7.5K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
7.5K
Calibration Curves: Correlation Coefficient
4.4K
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the...
4.4K
Kendall's Coefficient of Concordance
863
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
863
Extraction: Partition and Distribution Coefficients
4.5K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
4.5K
Wilcoxon Signed-Ranks Test for Matched Pairs
395
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
395