PromptFix: 逆のプロンプトチューニングによるバックドア削除
Tianrong Zhang1, Zhaohan Xi1, Ting Wang2
1School of Information Science & Technology, Pennsylvania State University.
まとめ
PromptFixは自然言語処理 (NLP) モデルのバックドアに対する新しい防御を提供します. この方法は,マルウェア型トリガートークンをモデルパラメータを変更することなく中和するために敵対的なプロンプトチューニングを使用して,少数のショット学習シナリオでセキュリティを強化します.
科学分野:
- 人工知能
- 自然言語処理
- 機械学習のセキュリティ
背景:
- 前もって訓練された言語モデル (PLM) は,優れたパフォーマンスを示しますが,特定のトリガートークンがモデル行動を操作するバックドアに脆弱です.
- PLMの汎用性と高いトレーニングコストのために,Few-Shotの微調整とプロンプトは人気のあるNLPトレーニングパラダイムです.
- 既存のバックドア対策には トリガー・インバーションと モデルの再訓練が必要で 効率が悪いこともあります
研究 の 目的:
- NLPモデルの新しいバックドア戦略であるPromptFixを導入します
- バックドア攻撃の脆弱性に対処する
- バックドアトリガーを効果的に無効化しながらモデルパラメータを保存する方法を開発する.
主な方法:
- PromptFixは,2組のソフトトークンを用いて対抗的なプロンプトチューニングを行います.一つはトリガーを近似し,もう一つはそれに抵抗します.
- この方法は,トリガーの明示的な逆転とモデルの微調整を避け,元のモデルのパラメータをそのまま保持します.
- トリガー識別とパフォーマンスの維持を適応的にバランスさせるため,敵対的な最適化が利用されます.
主要な成果:
- NLPモデルにおける様々なバックドア攻撃に対するPromptFixの有効性を実証しています.
- この方法は,ドメインシフト下でも強力な性能を示し,未知の予備トレーニングデータを持つモデルに適用できることを示しています.
- PromptFixは,モデルの全体的なパフォーマンスを損なうことなく,バックドアトリガーを成功裏に中和します.
結論:
- PromptFixは,いくつかのショット設定内のNLPモデルのバックドアを緩和するための有効でパラメータ効率の良いソリューションを提供します.
- この技術はドメインシフトに強固であり,現実のプロンプトチューニングアプリケーションに適しています.
- この対抗的なプロンプトチューニングアプローチは,事前訓練された言語モデルのセキュリティを強化するための有望な方向性を提供します.
関連する概念動画
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
3.3K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.3K
Enhanced Elimination of Poison
576
Poison can be effectively removed from the gastrointestinal (GI) tract through various decontamination procedures.
Antidotes serve a crucial role in counteracting the effects of poison by inhibiting enzymes responsible for producing harmful drug metabolites. In some cases, these toxic metabolites can be neutralized by endogenous cosubstrates, which are maintained at specific concentrations to prevent interaction with cellular macromolecules and subsequent cell death.
Renal excretion is the...
Antidotes serve a crucial role in counteracting the effects of poison by inhibiting enzymes responsible for producing harmful drug metabolites. In some cases, these toxic metabolites can be neutralized by endogenous cosubstrates, which are maintained at specific concentrations to prevent interaction with cellular macromolecules and subsequent cell death.
Renal excretion is the...
576
Masking and Demasking Agents
2.6K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.6K
Randomized Experiments
7.2K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
7.2K
Hindsight Biases
3.9K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.9K
Types of Errors: Detection and Minimization
2.3K
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
2.3K


