Related Experiment Video
Updated: Mar 28, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
SSA-KD: Self-structure-aware knowledge distillation for convolutional neural networks
Yiheng Lu1, Zhihui Zhang2, Ziyu Guan1
1Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University, No. 2 South TaiBai Road, Xian, 710071, China; School of Computer Science, Xidian University, No. 2 South TaiBai Road, Xian, 710071, China.
Abstract:
Knowledge Distillation has achieved great success in model compression for convolutional neural networks. However, the selection of the student model usually relies on universal small structures, which leads to plentiful incompatibility and inefficiency, i.e., the student model cannot be customized adaptively in terms of specified datasets and tasks. In this paper, we propose a self-structure-aware knowledge distillation to obtain the student model by personalizing the teacher model, namely, we first formulate a sub-network from the original teacher model, and then conduct the knowledge distillation across the teacher and student models. The customization of the student model is finalized via a structure-aware pruning method, which can yield a stable structure to ensure the effectiveness of the student model. Compared with previous knowledge distillation methods, our method can be executed with lower complexity while with the higher performance because the structure-aware pruning method can generate a layer-aligned sub-structure from the teacher model. It means the incompatibility and inefficiency concerns can be alleviated under the appropriate customization for the student model. We test the method on VGG-16, ResNet-32, and ResNet-50 with CIFAR-10 and CIFAR-100. Not only can we acquire the highest compression rate on the student model, but also require the lowest complexity to implement the method due to the same depth dimension between the teacher and the student model. Code is available at: https://github.com/motinwing/AFIE-distillation-SSA-KD.
Related Concept Videos
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
Convolution Properties I
The commutative property reveals that the input and the impulse response of an LTI (Linear Time-Invariant) system can be interchanged without affecting the output:
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Convolution: Math, Graphics, and Discrete Signals
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...