先行訓練された言語モデルの混乱の不一致性に基づくバックドアサンプル検出
Zuquan Peng1, Jianming Fu1, Lixin Zou1
1Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University, Wuhan, 430000, Hubei, China.
まとめ
私たちは新しいバックドアサンプル検出方法である Perturbation Discrepancy Consistency Evaluation (NETE) を導入し 毒性サンプルや膨大なリソースを必要とせずに 訓練済みモデルで悪意のあるデータを特定します NETEはトレーニングと推論の両段階でバックドア攻撃を効果的に検出します.
科学分野:
- 人工知能
- 機械学習のセキュリティ
背景:
- 前もって訓練されたモデルは 裏口の攻撃に脆弱です
- 既存の検出方法は,リソースやアクセス要件のため,しばしば非実用的です.
研究 の 目的:
- 実践的で効果的なバックドアサンプル検出方法を開発する.
- 訓練前の段階と訓練後の段階の両方で検出できるようにする.
主な方法:
- 混乱の不一致性評価 (NETE) を提案する.
- 既製の訓練済みのモデルと,混乱に対するマスク充填戦略を使用する.
- 一貫性を評価するために曲率を使用してログ確率の不一致を測定します.
主要な成果:
- NETEは,干渉の差異がクリーンなサンプルよりも少ないという現象を活用します.
- この方法は既存のゼロショットブラックボックス検出技術を上回ります.
- 4つの典型的なバックドア攻撃と5つの大きな言語モデルバックドア攻撃タイプに対して有効性を証明しました.
結論:
- NETEはバックドアサンプルを検出するための 実践的な解決策を提供します
- この方法は,さまざまな攻撃タイプとモデルフェーズで有効です.
- データ中毒に対する予め訓練されたモデルの安全性を向上させる.
さらに関連する動画
関連する概念動画
Survival Tree
159
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
159
Difference from Background: Limit of Detection
7.1K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
7.1K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K
Mismatch Repair
40.6K
Overview
40.6K
¹H NMR: Interpreting Distorted and Overlapping Signals
1.1K
Spin systems where the difference in chemical shifts of the coupled nuclei is greater than ten times J are called first-order spin systems. These nuclei are weakly coupled, and their chemical shifts and coupling constant can generally be estimated from the well-separated signals in the spectrum.
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
1.1K


