Related Experiment Video
Updated: Sep 21, 2026

Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
Distributed quantile regression for big data with missing covariates
Ye Fan1, Hua Zhou2,3, Jin Zhou3,4
1School of Statistics and Data Science, Capital University of Economics and Business, Beijing, 100070 China.
Abstract:
Data missing caused by non-response or drop-out appears routinely in modern medical studies. With the rapid development of information technology, medical data are becoming more and more massive in volume, and thus usually require distributed analysis in real-world applications, especially for cases where the data are collected from multiple centers or contain patient privacy. In this paper, we focus on developing efficient distributed algorithms to support quantile regression analysis in missing big data, which provides a powerful tool to model the skew and heterogeneous medical samples in reality, but currently remains a challenging issue. We employ the weighted quantile regression (WQR) technique to incorporate the missing information into the model and propose two two-stage distributed algorithms, IPW-ADMM and IPW-renewable, for efficient and privacy-preserving estimation of WQR. Both methods first estimate the missingness mechanism using a logistic model and then solve the WQR problem in a communication-efficient manner without sharing raw individual-level data. The IPW-ADMM algorithm parallelizes estimation using a multi-block alternating direction method of multipliers (ADMM), reformulating the nonsmooth WQR objective into a set of local subproblems. The IPW-renewable algorithm adopts a sequential renewable estimation framework with the smoothing technique, making it suitable for streaming or incremental data settings. Simulation studies demonstrate that both proposed methods achieve estimation accuracy comparable to the classical centralized interior point (IP) method, while offering substantial computational speedups in distributed environments. An application to the UK Biobank dataset further illustrates their practical utility in detecting heterogeneous genetic associations across quantiles under covariate missingness.
Supplementary Information:
The online version contains supplementary material available at https://doi.org/10.1007/s12561-026-09528-6.
Related Concept Videos
Distributions to Estimate Population Parameter
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the Guinness...
Quantifying and Rejecting Outliers: The Grubbs Test
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are observed.
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Statistical Methods for Analyzing Epidemiological Data