Related Experiment Video
Updated: Apr 5, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection
Caixia Yan1, Muyan Jiao1, Nuohan Xue1
1School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China.
This study introduces a new method for zero-shot object detection (ZSD) that synthesizes features for unseen classes. The Vision-Language Distillated Unseen Synthesizer (VLDUS) improves feature diversity and generalization by using knowledge distillation from CLIP.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Generative methods in zero-shot object detection (ZSD) synthesize features for unseen classes using semantic embeddings.
- Current methods suffer from poor diversity and generalization due to reliance on limited seen-class data for training feature synthesizers.
Purpose of the Study:
- To develop a novel knowledge distillation-based feature generation paradigm for ZSD.
- To enhance the diversity and generalization ability of synthesized unseen class features.
Main Methods:
- Proposed the Vision-Language Distillated Unseen Synthesizer (VLDUS) for ZSD.
- Implemented two complementary generative distillation strategies to transfer knowledge from a pre-trained CLIP model.
- Employed feature-aligned and relation-aligned generative distillation to improve generalization and intra-class diversity.
Main Results:
- VLDUS generates unseen features with high intra-class diversity and inter-class separability.
- Achieved significant performance improvements on ZSD and Generalized ZSD (GZSD) tasks.
- Outperformed state-of-the-art methods on benchmark datasets including MS COCO, PASCAL VOC, and DIOR.
Conclusions:
- VLDUS effectively addresses limitations in current ZSD generative methods.
- The proposed distillation strategies enhance the quality and applicability of synthesized features.
- VLDUS offers a robust solution for improved zero-shot object detection capabilities.
Related Concept Videos
Vision
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Depth Perception and Spatial Vision
Visual System
Once through the pupil, the light passes through the lens, a...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Light Acquisition
