Related Experiment Video
Updated: Sep 8, 2026

The Adventures of Fundi Intervention Based on the Cognitive and Emotional Processing in Attention Deficit Hyperactive Disorder Patients
Published on: June 12, 2020
Backdooring rationalization plus
Lei Wu1, Lingxiao Kong2, Fan He1
1Information Support Force Engineering University, Wuhan, China.
Abstract:
Rationalization models have recently garnered significant attention for enhancing the interpretability of natural language processing by first using a generator to select the most relevant pieces from the text with respect to the label, before passing the text input to the predictor. However, the robustness of the rationalization models is not sufficiently investigated. Specifically, this paper explores the robustness of rationalization models against backdoor attacks, which has been ignored by previous studies. Surprisingly, we find that conventional backdoor attack techniques fail to inject triggers into the rationalization model because its generator can filter out bad triggers. Considering this, we further propose a novel backdoor attack method named as BadRNL designed specially for the rationalization models. The core idea of BadRNL is first to search for the personalized trigger for each specific dataset and then manipulate the rationales and labels to conduct attacks. Besides, BadRNL controls the order of sample learning through poison-priority sampling strategies. Experimental results across five diverse NLP datasets-FEVER, MultiRC, Beer, Hotel, and Movie-demonstrate that BadRNL consistently achieves near 100% attack success rates (ASR). Crucially, the method maintains high classification accuracy and rationale quality on benign samples, highlighting a significant security risk for interpretable NLP systems deployed in real-world supply-chain scenarios.
Related Concept Videos
Hindsight Biases
Reason and Intuition
Rationalizing Substitutions
Rational Emotive Behavior Therapy
Impression Management Techniques III: Aligning Actions
Counterfactual Thinking