Related Experiment Video
Updated: Feb 13, 2026

Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology
Published on: March 8, 2024
Multiple interpretation ensemble distillation for graph neural networks
Kang Liu1, Yuqi Zhang1, Shunzhi Yang2
1School of Computer Science, South China Normal University, Guangzhou, 510000, China.
Abstract:
Existing graph knowledge distillation methods suffer from limited absorption of the teacher's "dark knowledge" because they rely on simple logit alignment, which often causes overfitting or incomplete capture of underlying patterns. Additionally, relying on a single perspective severely restricts the student's learning effectiveness and generalization ability. To address these issues, we develop a novel Multiple Interpretation Ensemble Distillation (MIED) method. It constructs a multi-interpreter composed of multiple single-layer MLPs for the student, termed the Student Interpretation (SI) component, to interpret knowledge from diversified outputs, thus avoiding representational bias from a single student output. Based on this, it introduces two effective strategies, i.e., Hybrid Sampling and Hierarchical Update. The former employs different sampling strategies for the outputs of the teacher and student (including the SI component). Specifically, the teacher's output adopts a percentage random sampler, while the outputs of the student and SI component both leverage a positive-negative sampler. With this design, MIED can facilitate better coordination of sample selection and the learning process among the teacher, student, and SI component. The latter updates the parameters of the last layer in the student using the exponential moving average of the fused parameters of the SI component, while the parameters of other layers are updated via a regular optimizer. This enhances the robustness and generalization performance of MIED. Extensive experiments on seven real-world public datasets demonstrate that MIED outperforms existing methods in node classification tasks, resulting in an average improvement of 5.56% over GCN and 27.43% over MLP, respectively. Moreover, compared with directly using multiple students (where the number is consistent with the number of layers in the SI component), MIED achieves improvements approximately 6.00% in time, 50.00% in space, and 0.20% in accuracy. These results indicate that MIED is scalable and generalizable, and exhibits robustness on complex samples.
Related Concept Videos
Multiple Bar Graph
Each bar or column in the multiple bar graph represents a data value. These graphs are used primarily in interrelating two or more sets of data. The categories of different kinds of data are listed along the horizontal or x-axis, whereas...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Distillation: Vapor–Liquid Equilibria
Ogive Graph
Graphing Antiderivatives
Graphs of Functions

