Distilling structural knowledge from CNNs to vision transformers for data-efficient visual recognition

Dingyao Chen1, Xiao Teng2, Xingyu Shen1

  • 1College of Computer Science and Technology, National University of Defense Technology, Changsha, 410073, Hunan, China.

Summary

This study introduces Feature-based Structural Knowledge Distillation (FSKD) to improve Vision Transformers (ViTs) by transferring CNN features. FSKD enhances ViT performance in visual recognition, especially with limited data.

Related Concept Videos

Vision01:24

Vision

Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
60.1K
Color Vision01:24

Color Vision

Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
1.5K
Distillation: Vapor–Liquid Equilibria01:01

Distillation: Vapor–Liquid Equilibria

Distillation is a separation technique that takes advantage of the boiling point properties of disparate elements in a mixture. To perform distillation, we begin by heating a miscible mixture of two liquids with a significant difference in boiling points (at least 20°C). As the solution heats up and reaches the bubble point of the more volatile component, some molecules of the more volatile component transition into the gas phase and travel upward into the condenser, which is a glass tube...
4.7K
Bacterial Transformation01:33

Bacterial Transformation

In 1928, bacteriologist Frederick Griffith worked on a vaccine for pneumonia, which is caused by Streptococcus pneumoniae bacteria. Griffith studied two pneumonia strains in mice: one pathogenic and one non-pathogenic. Only the pathogenic strain killed host mice.
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
60.1K
Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.0K
Transformers01:26

Transformers

A device that transforms voltages from one value to another using induction is called a transformer. A transformer consists of two separate coils, or windings, wrapped around the same soft iron core. However, they are electrically insulated from each other.
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
1.9K