Related Experiment Video
Updated: Aug 6, 2025

Infinium Assay for Large-scale SNP Genotyping Applications
Published on: November 19, 2013
Inferring feature importance with uncertainties with application to large genotype data
Pål Vegard Johnsen1,2, Inga Strümke3,4, Mette Langaas2
1SINTEF DIGITAL, Oslo, Norway.
Abstract:
Estimating feature importance, which is the contribution of a prediction or several predictions due to a feature, is an essential aspect of explaining data-based models. Besides explaining the model itself, an equally relevant question is which features are important in the underlying data generating process. We present a Shapley-value-based framework for inferring the importance of individual features, including uncertainty in the estimator. We build upon the recently published model-agnostic feature importance score of SAGE (Shapley additive global importance) and introduce Sub-SAGE. For tree-based models, it has the advantage that it can be estimated without computationally expensive resampling. We argue that for all model types the uncertainties in our Sub-SAGE estimator can be estimated using bootstrapping and demonstrate the approach for tree ensemble methods. The framework is exemplified on synthetic data as well as large genotype data for predicting feature importance with respect to obesity.
Related Concept Videos
Incomplete Dominance
Multiple Allele Traits
Evolutionary Relationships through Genome Comparisons
Polygenic Traits
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genetic Variation
Genes exist in different versions called alleles,...

