Related Experiment Videos
Bayesian nonparametric mixtures of categorical directed graphs for personalized causal inference
Federico Castelletti1, Laura Ferrini2
1Department of Statistical Sciences, Università Cattolica del Sacro Cuore, Largo Gemelli, 1,Milan 20123, Italy.
None:
Quantifying the causal effect of a treatment on a disease is a crucial task in medical science for the administration of effective therapies. Typically, such causal effects are inferred from multivariate data that are collected on patients and recorded in the form of categorical variables, including risk factors involved in disease progression, treatment assignments, and disease status. This feature motivates an approach to causal inference based on categorical Directed Acyclic Graphs (DAGs), which provide an effective framework for causal reasoning in complex multivariate settings. In this context, traditional DAG-based methods assume population homogeneity and accordingly attribute a unique causal effect to all subjects. However, this assumption is often unrealistic in clinical contexts, since patients may exhibit heterogeneous characteristics, possibly linked to unmeasured features. To address this issue, we propose a Bayesian nonparametric methodology based on a Dirichlet Process mixture of categorical DAGs, which allows treatment effects to vary across individuals because of underlying clustering structures in the data. We develop a Markov chain Monte Carlo algorithm for Bayesian posterior inference and evaluate our methodology through simulation studies. We then analyze patients affected by HER2+ breast cancer undergoing therapies that may cause cardiotoxic side effects. Importantly, our findings show that approaches neglecting population heterogeneity may produce biased results, since they can over- or under-estimate the risk of cardiotoxicity across patients.
Related Concept Videos
Causality in Epidemiology
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Bar Graph
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Multiple Bar Graph
Each bar or column in the multiple bar graph represents a data value. These graphs are used primarily in interrelating two or more sets of data. The categories of different kinds of data are listed along the horizontal or x-axis, whereas...