Related Experiment Video
Updated: Jan 23, 2026

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
1.3K
Overview and Practical Recommendations on Using Shapley Values for Identifying Predictive Biomarkers via CATE
David Svensson1, Erik Hermansson1,2, Nikolaos Nikolaou3
1AstraZeneca, Gothenburg, Sweden.
Statistics in Medicine
|January 22, 2026
Summary
This study introduces a novel surrogate estimation approach for Shapley Additive Explanations (SHAP) in Conditional Average Treatment Effect (CATE) modeling. This method efficiently identifies predictive biomarkers in high-dimensional data for precision medicine applications.
Area of Science:
- Machine Learning
- Causal Inference
- Explainable AI
Background:
- Individual Treatment Effect (ITE) modeling, specifically Conditional Average Treatment Effect (CATE) using meta-learners, is advancing causal inference from observational data.
- Explainable Machine Learning (XML), notably Shapley Additive Explanations (SHAP), enhances model interpretability in data science.
- The intersection of SHAP and CATE for predictive biomarker identification in precision medicine is underexplored.
Purpose of the Study:
- To address challenges in applying SHAP to multi-stage CATE strategies.
- To introduce a surrogate estimation approach for SHAP in CATE models.
- To enable efficient identification of predictive biomarkers using SHAP values in high-dimensional settings.
Main Methods:
- Developed a surrogate estimation approach for SHAP values in CATE modeling.
- The approach is agnostic to the specific CATE meta-learner strategy.
- Employed simulation benchmarking to evaluate biomarker identification accuracy.
Main Results:
- The proposed surrogate estimation method effectively reduces computational burden in high-dimensional data.
- Simulation results demonstrate accurate identification of biomarkers using SHAP values derived from various CATE meta-learners and Causal Forest.
- The approach facilitates the application of SHAP for biomarker discovery within CATE frameworks.
Conclusions:
- The surrogate SHAP estimation approach offers a computationally efficient and effective method for biomarker identification in CATE models.
- This work bridges the gap between explainable AI and causal inference for precision medicine.
- The findings support the use of SHAP for discovering predictive biomarkers in complex treatment effect modeling.
Related Concept Videos
Biodiversity and Human Values
16.4K
Human civilization relies on biodiversity in many ways. Sudden changes in species biodiversity result in environmental changes that can modify weather patterns and therefore human civilizations.
16.4K
Professional Values
10.4K
Nurses are responsible for caring for patients during birth, death, illness, and healing. Professional values guide the decisions and actions that nurses make in their careers. If nurses know the decisions and actions to take, providing patients with exceptional care is possible.
The values that are the foundation of the nursing profession are altruism, autonomy, human dignity, and social justice.
First, altruism refers to the concern for the welfare and well-being of others without personal...
The values that are the foundation of the nursing profession are altruism, autonomy, human dignity, and social justice.
First, altruism refers to the concern for the welfare and well-being of others without personal...
10.4K
Critical Values
10.2K
A critical value is a definite value obtained from a particular probability distribution at a predecided confidence level (or a predecided significance level) for a given population parameter. The critical value provides demarcation that separates the sample statistics that are likely to occur from the ones that are unlikely to occur based on the given probability distribution and the population parameter to be estimated. The critical value for normal distribution is obtained from the z...
10.2K
z Scores and Unusual Values
11.0K
The z score is one of the three measures of relative standing. It describes the location of a value in a dataset relative to the mean. z scores are obtained after the standardization of the values in a dataset. The z score for the mean is 0.
This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data...
This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data...
11.0K
Absolute and Local Extreme Values
56
The highest and lowest values of a function, relative to a reference axis, are known as extreme values. These include absolute maximum and absolute minimum values, which represent the highest and lowest points the function reaches across its entire domain. Within a restricted portion of the function, the highest and lowest values are referred to as local maximum and local minimum values, respectively.Periodic functions, such as sine and cosine, show extreme values at infinitely many points due...
56
Predicting Molecular Geometry
45.5K
VSEPR Theory for Determination of Electron Pair Geometries
45.5K

