Related Experiment Videos
Exploiting explanations for model extraction via knowledge distillation and mitigation with private counterfactuals.
Fatima Ezzeddine1,2, Silvia Giordano2, Omran Ayoub2
1Università della Svizzera italiana, Lugano, Switzerland.
Frontiers in Artificial Intelligence
|June 10, 2026
Summary
This study shows how explainable AI (XAI) can be exploited for model extraction attacks (MEA) in machine learning (ML) services. Differential privacy (DP) is explored as a defense to protect ML models and their data.
Area of Science:
- Artificial Intelligence
- Machine Learning Security
- Explainable AI (XAI)
Background:
- Machine learning models as a service (MLaaS) are increasingly deployed.
- Explainable AI (XAI) provides insights into ML model decisions, raising privacy concerns.
- Model explanations can be exploited for privacy leakage attacks like model extraction attacks (MEA).
Purpose of the Study:
- Investigate the exploitation of counterfactual explanations (CFs) for MEA in MLaaS platforms.
- Assess the effectiveness of differential privacy (DP) as a mitigation strategy against MEA.
- Propose a novel MEA approach leveraging CFs and Knowledge Distillation (KD).
Main Methods:
- Developed a novel MEA approach using Knowledge Distillation (KD) and counterfactual explanations (CFs).
- Proposed a method for training CF generators with integrated DP to produce private CFs.
- Conducted experiments on real-world datasets to evaluate the proposed MEA and DP mitigation.
Main Results:
- The proposed KD-based MEA successfully extracts high-fidelity substitute models with fewer queries than baselines.
- Integrating a privacy layer using DP effectively mitigates MEA.
- Balancing CF quality impacts explanation performance and MEA effectiveness.
Conclusions:
- Counterfactual explanations in MLaaS pose significant privacy risks exploitable via MEA.
- Differential privacy offers a viable mitigation strategy against explanation-based MEA.
- Further research is needed to balance privacy guarantees with the utility of explanations.
Related Concept Videos
Hindsight Biases
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
Counterfactual Thinking
Counterfactual thinking is a cognitive process wherein individuals mentally reconstruct alternative versions of past events, often beginning with “what if” or “if only.” This reflective mechanism plays a significant role in shaping emotional experiences and guiding future behavior. Though typically triggered by unfavorable or unexpected outcomes, counterfactual thinking can also emerge in mundane, everyday decisions and experiences, revealing its deep entrenchment in human cognition.Types of...
Fundamental Attribution Error
According to some social psychologists, people tend to overemphasize internal factors as explanations—or attributions—for the behavior of other people. They tend to assume that the behavior of another person is a trait of that person, and to underestimate the power of the situation on the behavior of others. They tend to fail to recognize when the behavior of another is due to situational variables, and thus to the person’s state. This erroneous assumption is called the fundamental attribution...
Theory of Attribution II: Kelley's Covariation Theory
Attribution theory plays a crucial role in social psychology, helping to explain how individuals interpret the causes of behavior. One prominent model within this field is Harold Kelley's covariation theory, which provides a systematic approach to determining whether internal traits or external circumstances drive a person's actions. The model posits that individuals rely on three key types of information—consensus, consistency, and distinctiveness—to make these judgments.Consensus: Comparing...
Theory of Attribution I: Correspondent Inference Theory
Correspondent inference theory, proposed by Jones and Davis in 1965, seeks to explain how individuals infer stable personality traits from observed behaviors. It suggests that people attribute actions to underlying dispositions rather than external circumstances, particularly when the behavior appears intentional and socially significant.Voluntary Behavior and Dispositional AttributionAccording to this theory, individuals are more likely to attribute behavior to personal traits when it appears...
Understanding Deception
Deception is a pervasive aspect of human communication. Empirical studies have shown that most individuals engage in some form of deceit on a daily basis, with approximately 20% of social exchanges involving deceptive elements. Lying follows a developmental trajectory, peaking during adolescence and declining with age, possibly due to the maturation of cognitive control and social accountability.Cognitive and Social Factors in Deception DetectionDespite its prevalence, accurately detecting...