基于预先训练的语言模型中的扰乱差异一致性的后门样本检测
Zuquan Peng1, Jianming Fu1, Lixin Zou1
1Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University, Wuhan, 430000, Hubei, China.
概括
我们推出了一种新的后门样本检测方法, 扰乱差异一致性评估 (NETE), 在训练和推断阶段,NETE有效地检测后门攻击.
科学领域:
- 人工智能
- 机器学习安全性
背景情况:
- 预先训练的模型很容易受到来自未经验证数据的后门攻击.
- 由于资源或访问要求,现有的检测方法往往不切实际.
研究的目的:
- 开发一种实用且有效的后门样本检测方法.
- 在培训前和培训后的阶段进行检测.
主要方法:
- 建议进行扰乱差异一致性评估 (NETE).
- 使用现成的预训练模型和对干扰的掩盖填充策略.
- 使用曲率测量日志概率差异以评估一致性.
主要成果:
- NETE利用了这种现象,比起干净的样本,后门样本的扰动差异变化较小.
- 这种方法优于现有的零射击黑盒检测技术.
- 对四种典型的后门攻击和五种大型语言模型后门攻击类型的有效性被证明.
结论:
- NETE提供了一种用于检测后门样本的实用解决方案.
- 这种方法在各种攻击类型和模型阶段都有效.
- 提升预训练模型的安全性,防止数据中毒.
更多相关视频
相关概念视频
Survival Tree
159
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
159
Difference from Background: Limit of Detection
7.1K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
7.1K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K
Mismatch Repair
40.6K
Overview
40.6K
¹H NMR: Interpreting Distorted and Overlapping Signals
1.1K
Spin systems where the difference in chemical shifts of the coupled nuclei is greater than ten times J are called first-order spin systems. These nuclei are weakly coupled, and their chemical shifts and coupling constant can generally be estimated from the well-separated signals in the spectrum.
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
1.1K


