Related Experiment Video
Updated: Sep 12, 2025

Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
Semi-parametric Bayes regression with network-valued covariates
Xin Ma1, Suprateek Kundu2, Jennifer Stevens3
1Department of Biostatistics and Bioinformatics, Emory University, Atlanta, USA.
None:
Although there has been an explosive rise in network data in a variety of disciplines, there is very limited development of regression modeling approaches based on high-dimensional networks. The scarce literature in this area typically assume linear relationships between the outcome and the high-dimensional network edges that results in an inflated model plagued by the curse of dimensionality and these models are unable to accommodate non-linear relationships or higher order interactions. In order to overcome these limitations, we develop a novel two-stage Bayesian non-parametric regression modeling framework using high-dimensional networks as covariates, which first finds a lower dimensional node-specific representation for the networks, and then embeds these representations in a flexible Gaussian process regression framework along with supplemental covariates for modeling the continuous outcome variable. Moving from edge-level analysis to node-level model allows us to scale up to high-dimensional networks, and enables node selection via an extension of the Gaussian process framework that involves spike-and-slab priors on the lengthscale parameters. Extensive simulations show a distinct advantage of the proposed approach in terms of prediction, coverage, and node selection. The proposed model achieves considerable gains when predicting posttraumatic stress disorder (PTSD) resilience based on brain networks in our motivating neuroimaging applications, and also identifies important brain regions associated with PTSD. In contrast, existing non-linear approaches that employ the full-edge set or those that use other dimension reduction techniques on the network are not equipped for node selection and results in poor prediction and characterization of predictive uncertainty, while linear approaches using the edge-level features are overly inflated and typically result in poor performance.
Related Concept Videos
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Distributions to Estimate Population Parameter
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

