Related Experiment Videos
Exploiting explanations for model extraction via knowledge distillation and mitigation with private counterfactuals
Fatima Ezzeddine1,2, Silvia Giordano2, Omran Ayoub2
1Università della Svizzera italiana, Lugano, Switzerland.
Abstract:
In recent years, there has been a notable increase in the deployment of machine learning (ML) models as services (MLaaS) across diverse production software applications. In parallel, explainable AI (XAI) continues to evolve, addressing the necessity for transparency in ML models. XAI techniques aim to enhance the transparency of ML models by providing insights, in terms of model's explanations, into their decision-making process. At the same time, some MLaaS platforms now offer explanations alongside the ML prediction outputs. This setup has elevated concerns regarding vulnerabilities in MLaaS, particularly in relation to privacy leakage attacks such as model extraction attacks (MEA). This is due to the fact that explanations can unveil insights about the inner workings of the model which could be exploited by malicious users. In this work, we focus on investigating how model explanations, particularly counterfactual explanations (CFs), can be exploited for performing MEA within the MLaaS platform. We also delve into assessing the effectiveness of incorporating differential privacy (DP) as a mitigation strategy. To this end, we first propose a novel MEA approach based on Knowledge Distillation (KD) that leverages CFs to effectively extract a substitute model of the target. Our approach operates without requiring any prior knowledge of the training data distribution by the attacker. Then, we advise an approach for training CF generators integrating DP to generate private CFs. We conduct thorough experimental evaluations on real-world datasets and demonstrate that our proposed KD-based MEA can yield a high-fidelity substitute model with a reduced number of queries with respect to baseline approaches. Furthermore, our findings reveal that including a privacy layer can allow mitigating the MEA. However, by balancing the quality of CFs, impacts the performance of the explanations and the MEA.
Related Concept Videos
Hindsight Biases
Counterfactual Thinking
Fundamental Attribution Error
Theory of Attribution II: Kelley's Covariation Theory
Theory of Attribution I: Correspondent Inference Theory
Understanding Deception