Related Experiment Video
Updated: May 29, 2026

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
Nonlinear kernel-based high-dimensional inference for set-based genetic association studies
Zechen Zhang1,2,3, Hui Yang1,2,3, Meilin Zhu1
1Division of Health Statistics, School of Public Health, Hebei Medical University, 361 East Zhongshan Road, Shijiazhuang, Hebei 050017, P.R. China.
Abstract:
Nonlinear genetic architectures, including epistasis and threshold effects, are increasingly recognized as contributors to complex disease risk, yet most existing SNP-set association tests rely on linear modeling assumptions, resulting in reduced power and unstable inference when genetic effects are nonlinear or heterogeneously distributed across variants. To address this limitation, we propose a nonlinear high-dimensional inference framework for set-based genetic association analysis that integrates scalable kernel representations with valid statistical inference. The framework combines distance correlation-based sure independence screening to reduce ultra-high dimensional predictors, kernel principal component analysis with Nyström approximation for nonlinear feature extraction, and de-sparsified LASSO to enable asymptotically valid hypothesis testing in high dimensions, together with a two-stage omnibus testing strategy that adaptively aggregates evidence across complementary signal models. Extensive simulation studies demonstrate that the proposed method maintains well-calibrated Type I error and consistently achieves higher power than established set-based approaches, including Sequence Kernel Association Test and adaptive Sum of Powered Score test, particularly under nonlinear and heterogeneous genetic effect scenarios, while remaining competitive in linear settings. Application to Alzheimer's Disease Neuroimaging Initiative data identifies gene-level associations with brain regional volumes that converge on neuronal excitability, calcium signaling, and cytoskeletal regulation, biological processes centrally implicated in neurodegeneration. Together, this work provides a robust and scalable framework for nonlinear set-based inference in genome-wide studies, expanding the analytical toolbox for dissecting complex genetic contributions to disease.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Genetic Variation
Genes exist in different versions called alleles, which...
Multiple Allele Traits
Multiple Allele Traits
