Related Experiment Video
Updated: Aug 10, 2025

Author Spotlight: Analgesic Effect of Tuina on Rat Models with Compression of the Dorsal Root Ganglion Pain
Published on: July 14, 2023
LAD: Layer-Wise Adaptive Distillation for BERT Model Compression
Ying-Jia Lin1, Kuan-Yu Chen1, Hung-Yu Kao1
1Department of Computer Science and Information Engineering, National Cheng Kung University, Tainan 70101, Taiwan.
Abstract:
Recent advances with large-scale pre-trained language models (e.g., BERT) have brought significant potential to natural language processing. However, the large model size hinders their use in IoT and edge devices. Several studies have utilized task-specific knowledge distillation to compress the pre-trained language models. However, to reduce the number of layers in a large model, a sound strategy for distilling knowledge to a student model with fewer layers than the teacher model is lacking. In this work, we present Layer-wise Adaptive Distillation (LAD), a task-specific distillation framework that can be used to reduce the model size of BERT. We design an iterative aggregation mechanism with multiple gate blocks in LAD to adaptively distill layer-wise internal knowledge from the teacher model to the student model. The proposed method enables an effective knowledge transfer process for a student model, without skipping any teacher layers. The experimental results show that both the six-layer and four-layer LAD student models outperform previous task-specific distillation approaches during GLUE tasks.
Related Concept Videos
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
Survival Tree
Building a Survival Tree
Constructing a...
Improving Translational Accuracy
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Compartment Models: Single-Compartment Model
Distillation: Vapor–Liquid Equilibria
