Related Experiment Video
Updated: Sep 9, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Adaptive metric for knowledge distillation by deep Bregman divergence
Tongtong Yuan1, Zixuan Xu2, Bo Liu1
1Beijing University of Technology, China.
Abstract:
Knowledge distillation (KD) makes it possible to deploy high-accuracy models on devices with limited resources and is an effective means of achieving lightweight models. With the advancement of technology, the methods of knowledge distillation are also continuously developing and improving to adapt to different application scenarios and needs. To facilitate the transfer of knowledge from larger networks to smaller and lighter networks, KD has been employed to bridge the gap in probability outputs or middle-layer representations between teacher and student networks. Unlike the consistent probability outputs observed between teacher and student networks, the middle-layer representations exhibit significant variations in both structure and distribution. Traditional metrics such as Euclidean distance or MSE treat all intermediate features uniformly and do not adapt to the heterogeneous characteristics of feature distributions across different layers or models. These fixed metrics often fail to account for the spatial, semantic, and statistical variations between teacher and student networks. To address this limitation, we propose using a parameterized and adaptive metric based on deep Bregman divergence. This divergence function is learned from data, enabling the measurement to adjust to the underlying feature distributions at different layers, leading to more effective and robust knowledge transfer. Importantly, our method can also serve as a complementary enhancement (i.e., x+Bregman) to almost all other KD methods focused on distilling probability outputs. Extensive experiments demonstrate that our approach outperforms many existing KD methods, achieving superior performance across diverse datasets and network models.
Related Concept Videos
Mean Absolute Deviation
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Divergence and Stokes' Theorems
Improving Translational Accuracy
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Dot Product: Problem Solving
Identify the problem: Start by reading the problem and...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
