Related Experiment Video
Updated: Mar 18, 2026

Oncogenic Gene Fusion Detection Using Anchored Multiplex Polymerase Chain Reaction Followed by Next Generation Sequencing
Published on: July 5, 2019
Causal inference and the data-fusion problem
Elias Bareinboim1, Judea Pearl2
1Department of Computer Science, University of California, Los Angeles, CA 90095; Department of Computer Science, Purdue University, West Lafayette, IN 47907 eb@purdue.edu.
Abstract:
We review concepts, principles, and tools that unify current approaches to causal analysis and attend to new challenges presented by big data. In particular, we address the problem of data fusion-piecing together multiple datasets collected under heterogeneous conditions (i.e., different populations, regimes, and sampling methods) to obtain valid answers to queries of interest. The availability of multiple heterogeneous datasets presents new opportunities to big data analysts, because the knowledge that can be acquired from combined data would not be possible from any individual source alone. However, the biases that emerge in heterogeneous environments require new analytical tools. Some of these biases, including confounding, sampling selection, and cross-population biases, have been addressed in isolation, largely in restricted parametric models. We here present a general, nonparametric framework for handling these biases and, ultimately, a theoretical solution to the problem of data fusion in causal inference tasks.
More Related Videos
Related Concept Videos
Cause and Effect
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Criteria for Causality: Bradford Hill Criteria - II
Criteria for Causality: Bradford Hill Criteria - I
Correlation and Causation
Correlation versus Causation
If the dependent variable increases or decreases when the independent variable increases, there is a positive or negative...
Causality in Epidemiology

