Related Experiment Video
Updated: Jan 7, 2026

Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
Informative Co-Data Learning for High-Dimensional Horseshoe Regression
Claudio Busatto1, Mark A van de Wiel2
1Department of Statistics, Computer Science, Applications "G. Parenti,", University of Florence, Florence, Italy.
We introduce informative Horseshoe regression (infHS), a Bayesian model that improves high-dimensional regression by incorporating prior knowledge (co-data). This method enhances variable selection and prediction accuracy in genomics.
Area of Science:
- Genomics
- Biostatistics
- Computational Biology
Background:
- High-dimensional data are common in clinical genomics for identifying trait predictors.
- Incorporating prior knowledge (co-data) can enhance predictive model performance.
Purpose of the Study:
- To develop a novel Bayesian regression model for high-dimensional data that integrates co-data.
- To improve variable selection and prediction accuracy by leveraging external information.
Main Methods:
- Developed the informative Horseshoe regression (infHS) model.
- Implemented Gibbs sampler for moderate dimensions and Variational approximation for large-scale data.
- Regressed prior variances of regression parameters on co-data variables.
Main Results:
- The simulation study demonstrated the benefits of including co-data.
- The infHS model showed superior performance compared to existing methods in two genomics applications.
- The Variational approximation enabled efficient analysis of very large datasets.
Conclusions:
- The infHS model effectively incorporates co-data into high-dimensional regression.
- This approach offers improved variable selection and predictive performance in genomics.
- infHS provides flexible computational tools for different data scales and inference goals.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Correlation and Regression
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
