Related Experiment Videos
Understanding and Optimizing Counterfactual Generation in LLMs via Contribution Attribution and Target Intervention
Abstract:
Counterfactual generation in large language models (LLMs), designed to enhance the comprehensiveness of semantic information in datasets, is an important technique for data augmentation. While existing research has evaluated counterfactual generation across various scenarios, there is a lack of further exploration into the underlying mechanisms, leading to an inability to effectively control the quality of the counterfactual generation. To fill this gap, we conducted mechanistic interpretability research on the task of sentiment counterfactual generation and proposed the contribution mechanisms, including task learning and task execution. During the mechanistic investigation, we identified the multihead attention (MHA) module as the critical component influencing counterfactual generation and pinpointed specific attention heads as the functional units of the contribution mechanisms. Finally, building upon the comprehensive mechanistic framework, we propose an optimization strategy achieved through targeted interventions on key components. Experiments on the Stanford Sentiment Treebank (SST) dataset, evaluated via an automated quality metric, demonstrate that our method yields consistent and substantial improvements. In particular, the generation quality scores are increased by +0.35 (from 6.75 to 7.10) on Llama2-7B, +1.06 (from 3.80 to 4.86) on Llama2-13B, and +0.36 (from 5.53 to 5.89) on Baichuan-13B.
Related Concept Videos
Theory of Attribution I: Correspondent Inference Theory
Attribution Theory
Counterfactual Thinking
Theory of Attribution II: Kelley's Covariation Theory
Fundamental Attribution Error
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...