Related Experiment Video
Updated: Feb 9, 2026

Using Looming Visual Stimuli to Evaluate Mouse Vision
Published on: June 13, 2019
Distilling structural knowledge from CNNs to vision transformers for data-efficient visual recognition
Dingyao Chen1, Xiao Teng2, Xingyu Shen1
1College of Computer Science and Technology, National University of Defense Technology, Changsha, 410073, Hunan, China.
This study introduces Feature-based Structural Knowledge Distillation (FSKD) to improve Vision Transformers (ViTs) by transferring CNN features. FSKD enhances ViT performance in visual recognition, especially with limited data.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Knowledge distillation (KD) transfers model representations, typically aligning output logits.
- Existing methods for CNN-to-ViT transfer overlook rich semantic structures in CNN features.
- This limits Vision Transformers (ViTs) in inheriting convolutional neural network (CNN) inductive biases.
Purpose of the Study:
- Propose a Feature-based CNN-to-ViT Structural Knowledge Distillation (FSKD) framework.
- Integrate semantic structural knowledge from CNN features with ViT's long-range dependency capabilities.
- Enhance ViT performance in visual recognition, particularly in low-data regimes.
Main Methods:
- Develop a feature alignment module to bridge CNN and ViT representational gaps.
- Incorporate a global feature alignment loss.
- Introduce patch-wise and attention-wise distillation losses for inter-patch similarity and attention distribution transfer.
Main Results:
- FSKD effectively transfers semantic structural knowledge from CNNs to ViTs.
- The framework significantly improves ViT performance in visual recognition tasks.
- Performance gains are particularly notable in scenarios with limited training data.
Conclusions:
- FSKD offers a novel approach to knowledge distillation from CNNs to ViTs.
- The method successfully transfers rich structural information beyond simple logit alignment.
- FSKD demonstrates the potential for improved ViT generalization and efficiency, especially in data-scarce environments.
More Related Videos
09:29A Standardized Obstacle Course for Assessment of Visual Function in Ultra Low Vision and Artificial Vision
Published on: February 11, 2014
10:23Author Spotlight: A Machine-Vision Approach to Transmission Electron Microscopy Workflows, Results Analysis and Data Management
Published on: June 23, 2023
Related Concept Videos
Vision
Color Vision
Distillation: Vapor–Liquid Equilibria
Bacterial Transformation
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
Depth Perception and Spatial Vision
Transformers
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...