Related Experiment Video
Updated: Jan 22, 2026

P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
Published on: September 8, 2023
When Not to Classify: Anomaly Detection of Attacks (ADA) on DNN Classifiers at Test Time
David Miller1, Yujia Wang2, George Kesidis3
1School of EECS, Penn State, University Park, PA, 16802, U.S.A. djmiller@engr.psu.edu.
Adversarial attacks threaten machine learning systems. This study introduces an unsupervised anomaly detector that effectively identifies these attacks, even subtle ones, outperforming existing methods for enhanced security.
Area of Science:
- Machine Learning Security
- Deep Neural Networks
- Adversarial Attacks
Background:
- Machine learning systems, particularly deep neural networks (DNNs), face significant threats from adversarial learning attacks, specifically evasion attacks at test time.
- Existing defenses often focus on robustifying classifiers against perturbed inputs, but this is not always optimal or feasible.
- In certain scenarios, detecting an attack is more valuable than correctly classifying a perturbed input, especially when the attacker is the sole recipient of the decision.
Purpose of the Study:
- To propose and evaluate a novel unsupervised anomaly detection (AD) method for identifying adversarial perturbations in deep neural networks.
- To demonstrate that adversarial perturbations are machine detectable, even when small.
- To offer a detection-based defense strategy that is actionable in specific adversarial scenarios.
Main Methods:
- Developed a purely unsupervised anomaly detector (AD) that models joint densities of deep layers using null hypothesis density models.
- Incorporated multiple DNN layers, source/destination class concepts, class confusion matrices, and DNN weight information.
- Constructed a novel decision statistic based on Kullback-Leibler divergence for attack detection.
Main Results:
- The proposed AD method outperformed previous detection methods on MNIST and CIFAR image databases under three attack strategies.
- Achieved strong Receiver Operating Characteristic (ROC) area under the curve (AUC) detection accuracy for two attacks.
- Demonstrated superior accuracy compared to recent methods on the strongest (CW) attack and effectiveness against white-box reverse engineering attacks.
Conclusions:
- Adversarial perturbations are detectable using unsupervised anomaly detection, even subtle ones.
- The novel AD method provides a robust and effective defense against various adversarial evasion attacks.
- The approach offers a valuable alternative or complement to classifier robustification, particularly in scenarios where attack detection is paramount.
More Related Videos
08:14MicroRNA Based Liquid Biopsy: The Experience of the Plasma miRNA Signature Classifier MSC for Lung Cancer Screening
Published on: October 26, 2017
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Related Concept Videos
Classifying Matter by Composition
According to its composition, the matter can be classified into two broad categories — pure substances and mixtures.
A pure substance is a form of matter that has a constant composition throughout with uniform properties. For example, any sample of sucrose has the same composition and same physical properties, such as melting point, color, and sweetness, regardless of the source from which it is isolated.
A mixture is composed of two or...
Classifying Matter by State
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Nursing Interventions II: Selecting and Classifying the Nursing Interventions
Acid Attack on Concrete
The rate at which hydrogen...