Related Experiment Video
Updated: Jan 11, 2026

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Multi-teacher knowledge distillation framework for lightweight anomaly detection
Behnam Yousefimehr1, Mehdi Ghatee1, Roozbeh Razavi-Far2
1Department of Mathematics and Computer Science, Amirkabir University of Technology, Hafez Ave., Tehran, 15875-4413, Tehran, Iran.
Abstract:
Anomaly detection is essential in various domains, where identifying rare and irregular samples is critical for system safety and security. This problem encounters a significant challenge due to extreme class imbalance, where normal samples vastly outnumber anomalous ones. This disparity poses difficulties for traditional learning models in effectively identifying anomalies. This paper introduces a novel framework that, for the first time, integrates knowledge distillation with multiple resampling strategies to address imbalanced learning while incorporating model compression for efficient deployment. The proposed method trains multiple teacher models on datasets resampled using diverse oversampling and undersampling techniques. By distilling knowledge from these teachers, the student model learns a balanced representation of normal and anomalous samples while maintaining a compact structure. Additionally, this paper provides a theoretical analysis showing that the proposed knowledge distillation algorithm correctly identifies class distinctions. This algorithm enhances generalization, reduces overfitting, and improves robustness in the presence of corrupted or noisy data, thereby demonstrating its practical utility in diverse and challenging conditions. Although the training process requires additional computational resources due to the multi-teacher setup, the resulting compressed student model offers significant advantages in terms of accuracy, efficiency, and inference speed, making it highly suitable for real-time anomaly detection applications. Furthermore, we have evaluated the proposed MTKD framework across six datasets, including WUSTL-EHMS, Credit-Card-Fraud, TON-IoT, KDD99, HYPERAKTIV, and ICU-IoT-Flock, covering domains such as fraud detection, intrusion detection, and healthcare monitoring, which demonstrates its domain-agnostic effectiveness in diverse real-world scenarios.
Related Concept Videos
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Mass Analyzers: Overview
Detection of Black Holes
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Mass Analyzers: Common Types
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...