Related Experiment Video
Updated: Nov 7, 2025

A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Memory-Replay Knowledge Distillation
Jiyue Wang1, Pei Zhang2, Yanxiong Li1
1School of Electronic and Information Engineering, South China University of Technology, Guangzhou 510641, China.
Memory-replay Knowledge Distillation (MrKD) uses historical models as teachers for improved deep neural network training. This self-knowledge distillation method stabilizes learning by regularizing with past outputs, enhancing performance across image and audio datasets.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Computer Vision
- Signal Processing
Background:
- Knowledge Distillation (KD) is crucial for Deep Neural Network (DNN) compression in intelligent sensor systems.
- Self-Knowledge Distillation (self-KD) methods often require task-specific DNN redesign or effective data augmentation.
- Existing self-KD approaches have limitations hindering widespread adoption.
Purpose of the Study:
- To introduce a novel self-KD method, Memory-replay Knowledge Distillation (MrKD), overcoming limitations of prior approaches.
- To leverage historical models as teachers within the self-KD framework.
- To enhance DNN training stability and performance without external knowledge dependencies.
Main Methods:
- Proposed a self-KD training strategy penalizing KL divergence between current and historical model outputs.
- Utilized a Fully Connected Network (FCN) to ensemble historical teacher outputs for guidance.
- Implemented Knowledge Adjustment (KA) to correct teacher logit outputs for ground truth accuracy.
Main Results:
- MrKD demonstrated improved single model training efficiency and performance.
- The method proved effective across diverse image (CIFAR-100, CIFAR-10, CINIC-10) and audio (DCASE) datasets.
- MrKD highlighted the value of utilizing historical models in DNN training.
Conclusions:
- MrKD offers an effective and practical self-KD approach by utilizing historical models.
- The method provides a stable learning regularization strategy through historical output distributions.
- MrKD presents a promising direction for DNN compression and performance enhancement.
More Related Videos
07:26The Deese-Roediger-McDermott DRM Task: A Simple Cognitive Paradigm to Investigate False Memories in the Laboratory
Published on: January 31, 2017
08:53Using a Classroom-Based Deese Roediger McDermott Paradigm to Assess the Effects of Imagery on False Memories
Published on: November 14, 2018
Related Concept Videos
Explicit Memories
Episodic memory contains information about personally experienced events and is reported as a story. An example of episodic memory is recalling a birthday celebration. This type of memory includes the what, where, and when of an event, as...
Interference and Decay
Interference occurs when competing memories hinder the retrieval of particular information. It can be classified into two types: proactive and retroactive interference. Proactive...
Understanding Memory
Chunking and Rehearsal in Sensory Memory
Long-Term Memory
Long-term memory can be categorized into two primary types: explicit and implicit memory. Explicit memory, also known as declarative memory, involves the conscious recollection of information that we deliberately try to remember, recall, and articulate. This type of memory encompasses specific facts, events, and...
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...