Related Experiment Video
Updated: Jul 17, 2026

13:19
Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
Unified knowledge distillation: Integrating offline, online, and self-distillation paradigms for image classification
Shuang Wang1, Zhenye Chen1, Fuli Wu1
1Zhejiang University of Technology, Liuhe Road No.288, Hangzhou, Zhejiang, China.
Summary
Unified Knowledge Distillation (UKD) effectively integrates offline, online, and self-knowledge distillation paradigms. This framework improves model adaptability and efficiency across diverse architectures by using a proxy teacher and dual-teacher collaboration.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Deep Learning
Background:
- Knowledge distillation (KD) has three main paradigms: offline, online, and self-KD.
- These paradigms face limitations in adaptability, stability, and expressiveness.
- Integrating these paradigms is challenging due to differing supervision and complexity.
Purpose of the Study:
- To propose a Unified Knowledge Distillation (UKD) framework.
- To efficiently transfer knowledge across diverse architectures and model capacities.
- To address limitations of individual KD paradigms and their integration challenges.
Main Methods:
- UKD coordinates three heterogeneous KD paradigms in a single optimization process.
- A lightweight proxy teacher bridges paradigms and reduces capacity gaps.
- A dual-teacher collaboration strategy dynamically adjusts teaching intensity.
Main Results:
- UKD demonstrates effectiveness and efficiency across CNN, ViT, and MLP architectures.
- Achieved performance gains of up to 11.38% on CIFAR-100 and 2.68% on ImageNet-1K.
- Showcased significant improvements with modest computational overhead.
Conclusions:
- UKD offers a unified approach to knowledge distillation.
- The framework enhances model performance and adaptability.
- UKD provides an efficient solution for cross-architecture knowledge transfer.
