Related Experiment Video
Updated: Oct 9, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
A dual-level sparsity Bayesian framework for rare-variant association analysis using integrated nested laplace
1Department of Mathematics and Computer Science, Hetao College, Bayannur 015000, P. R. China.
Abstract:
Background. Rare genetic variants (minor allele frequency, MAF <1%) carry a substantial share of the unexplained heritability of complex human traits, yet set-based association tests lose power precisely in the regime in which rare variation is most informative - when only a small minority of the variants inside a gene is causal and the aggregate contribution to phenotypic variance is well below one per cent. Burden tests dilute a genuine signal across neutral polymorphisms because they impose a common effect direction, whereas variance-component tests such as the optimal unified sequence kernel association test (SKAT-O) estimate a single dispersion parameter shared by every variant in the set and therefore cannot isolate the few variants that actually drive the association. Methods. We present a Bayesian INLA Model (BIM), a hierarchical Bayesian framework that couples the Integrated Nested Laplace Approximation (INLA) with a dual-level sparsity prior. At the gene level, a bimodal mixture prior on the log-precision of the gene-specific variance component performs explicit model selection between an associated and a null state. At the variant level, a horseshoe prior supplies adaptive, variant-specific shrinkage whose posterior shrinkage factor admits a closed-form characterization, so that individual pathogenic variants escape penalization while neutral effects are driven towards zero. Functional annotations enter as hierarchical covariates that modulate both the location and the scale of the variant-effect prior. Gene-level evidence is summarized by marginal-likelihood Bayes factors and posterior inclusion probabilities, and discoveries are declared by a Bayesian false-discovery-rate (FDR) rule that controls the average local FDR of the selected set. Posterior uncertainty is decomposed into parametric, structural, internal (genotype uncertainty) and external (population structure) components through an explicit application of the law of total variance. Results. Across 100 replicated simulations calibrated on 1000 Genomes Project Phase 3 European haplotypes, BIM attained 75.2% power (95% confidence interval [CI] 69.5-81.0%) to detect causal genes in the most demanding scenario of 0.5% variance explained, against 58.0% for BATI, 38.0% for MiST, 22.0% (95% CI 16.9-27.1%) for SKAT-O and 9.0% (95% CI 5.4-12.6%) for the burden test (paired t-test [Formula: see text] for every pairwise comparison). The realized FDR was 3.8%, below the nominal 5% level, and gene-level posterior inclusion probabilities were well calibrated against the empirical frequency of true association. In a whole-exome sequencing analysis of 500 chronic lymphocytic leukemia (CLL) cases and 1300 ancestry-matched controls, BIM returned a Bayesian FDR 5% discovery set of 12 genes, including the established susceptibility genes BRCA2 (Bayes factor, BF [Formula: see text] 50.3) and CHEK2 (BF [Formula: see text] 30.1) and the novel candidate ABCD3 (BF [Formula: see text] 22.4), a peroxisomal ABC transporter with emerging links to cancer metabolism. The genomic inflation factor was [Formula: see text], and INLA posteriors agreed with gold-standard Hamiltonian Monte Carlo to a Pearson correlation of 0.9987 at a small fraction of the computational cost. Conclusion. Placing sparsity at two biological scales simultaneously, and solving the resulting latent Gaussian model with INLA rather than Markov chain Monte Carlo, converts a computationally prohibitive Bayesian formulation into a practical genome-scale tool. BIM delivers three- to four-fold power gains over SKAT-O in the sparse architectures characteristic of complex-trait genetics while retaining variant-level interpretability and calibrated error control. The open-source implementation is available at https://github.com/meibujun/BIM-INLA.
Related Concept Videos
Introduction To Survival Analysis
The primary goal of survival analysis is to estimate survival time—the time until a...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Assumptions of Survival Analysis
Single Nucleotide Polymorphisms-SNPs
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Friedman Two-way Analysis of Variance by Ranks
