Related Experiment Video
Updated: Jul 8, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Debiased inference for heterogeneous subpopulations in a high-dimensional logistic regression model
Hyunjin Kim1, Eun Ryung Lee2, Seyoung Park3
1Department of Statistics, Sungkyunkwan University, Seoul, 100190, South Korea.
This study introduces a new statistical method to analyze complex, heterogeneous data in cancer cell lines. The fused group Lasso approach effectively quantifies covariate effects across subpopulations, improving inference for high-dimensional binary responses.
Area of Science:
- Statistics
- Bioinformatics
- Computational Biology
Background:
- Data heterogeneity is common in scientific studies, particularly with complex datasets and subpopulations.
- Analyzing high-dimensional binary responses in heterogeneous data presents significant inferential challenges.
- Existing methods often struggle to effectively quantify covariate effects across subpopulations.
Purpose of the Study:
- To develop a novel statistical inference method for high-dimensional logistic regression in the presence of heterogeneous subpopulations.
- To investigate and quantify heterogeneity in covariate effects across subpopulations.
- To test the equivalence and significance of covariate effects in complex datasets.
Main Methods:
- Proposed a fused group Lasso penalization method for overall sparsity and coefficient fusion.
- Developed a bias-corrected statistical inference method.
- Adapted proximal gradient and alternating direction method of multipliers (ADMM) for computational efficiency.
- Provided non-asymptotic analyses for the fused group Lasso and chi-squared approximations for debiased test statistics.
Main Results:
- The proposed method effectively accounts for data heterogeneity in high-dimensional logistic regression.
- Simulations demonstrated superior performance compared to existing methods.
- The fused group Lasso achieved sparsity and fusion of coefficients across subpopulations.
- Debiased test statistics were shown to admit chi-squared approximations.
Conclusions:
- The novel statistical inference method offers a robust approach for analyzing heterogeneous subpopulations with high-dimensional binary data.
- The method provides accurate and efficient statistical inference, outperforming existing techniques.
- Demonstrated practical utility through analysis of Cancer Cell Line Encyclopedia (CCLE) data.
More Related Videos
Related Concept Videos
Test for Homogeneity
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Distributions to Estimate Population Parameter
Expected Frequencies in Goodness-of-Fit Tests
Choosing Between z and t Distribution
Goodness-of-Fit Test

