Related Experiment Video
Updated: May 5, 2026

Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
CIRCUS: A Causal Intervention-Based Framework for Enhancing Counterfactual Fairness in Trained Classifiers
None:
Ensuring model fairness for preventing potential biases based on any sensitive attribute is crucial for the societal acceptance of artificial intelligence in critical applications. Among various fairness concepts, counterfactual fairness has gained prominence as it is grounded in causal inference. This concept requires that an individual's prediction in the original world remains consistent with that in the counterfactual world where the sensitive feature value is modified. In this article, we aim to mitigate counterfactual biases of the model through causal intervention. Specifically, we first achieve effective causal intervention and counterfactual generation by proposing the causal inference tabular generative adversarial network (CITGAN) architecture. Unlike prior approaches based on variational autoencoders (VAEs) that inherently lack structural causal model (SCM) fidelity due to simultaneous generation, CITGAN strictly enforces causal consistency via an end-to-end topological generation process. By integrating exogenous variable inference with sequential generation, CITGAN ensures that functional dependencies are structurally preserved. Building on CITGAN, we propose the CIRCUS framework, a causal intervention-based framework designed to intuitively enhance the counterfactual fairness in trained classifiers. CIRCUS generates counterfactually discriminatory samples (CDSs) via causal intervention, guided by gradients and feature contributions, and subsequently applies bias correction preprocessing to their labels for classifier retraining. Experimental results demonstrate that CIRCUS effectively enhances counterfactual fairness while maintaining robust classification performance. Specifically, for the deep neural network (DNN) model, the $\text {MMD}_{\text {L}}$ and $\text {MMD}_{\text {K}}$ values are reduced by averages of 39.7% and 40.4%, respectively, compared with the second-best result. For the residual network (ResNet) model, these reductions amount to 56.7% and 54.5%, respectively.
Related Concept Videos
Cause and Effect
Hindsight Biases
Causality in Epidemiology
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Counterfactual Thinking
Halo Effect

