Related Experiment Video
Updated: Jan 11, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Robust learning for ridge-penalized quasi-GLMs under non-identical distributions.
Huiming Zhang1, Wan Tian2, Qiuran Yao3
1Institute of Artificial Intelligence, Beihang University, Beijing, China.
This study introduces a robust log-truncated minimization estimator and stochastic gradient descent (SGD) algorithm for quasi-generalized linear models, offering improved performance with outlier-prone data.
Area of Science:
- Statistics and Machine Learning
- Robust Statistics
- Optimization
Background:
- Robust estimation is crucial for statistical and machine learning models dealing with outliers and inverse problems.
- Existing methods often rely on assumptions of light-tailed error distributions, limiting their applicability.
- There is a need for robust estimators that can handle independent non-identical distributed (i.n.i.d.) data without stringent moment assumptions.
Purpose of the Study:
- To introduce a novel log-truncated minimization estimator for quasi-generalized linear models.
- To develop a corresponding stochastic gradient descent (SGD) algorithm for efficient optimization.
- To provide a robust alternative to ordinary generalized linear models (GLMs) that accommodates data with outliers and i.n.i.d. sampling.
Main Methods:
- Development of a log-truncated minimization estimator.
- Derivation of non-asymptotic excess risk and [Formula: see text]-risk bounds for i.n.i.d. data using log-truncated Lipschitz losses.
- Analysis of the iteration complexity of SGD for non-convex log-truncated minimization.
- Empirical validation using simulated and real-world datasets, including German health care demand data.
Main Results:
- The proposed log-truncated minimization estimator offers robustness without requiring light-tailed error distributions or finite variance.
- Non-asymptotic risk bounds are derived under a finite β-th moment assumption for the data.
- The developed SGD algorithm demonstrates superior performance compared to non-robust methods in empirical evaluations.
- The practical utility is confirmed through a robust negative binomial regression analysis on health care data.
Conclusions:
- The novel log-truncated minimization approach effectively robustifies objective functions for quasi-generalized linear models.
- The associated SGD algorithm is efficient and performs well empirically, even with non-convex objectives.
- This method provides a valuable tool for analyzing complex datasets with outliers and i.n.i.d. properties.
Related Concept Videos
Quadratic Models
Distributions to Estimate Population Parameter
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...

