Related Experiment Video
Updated: Jun 13, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Global and Local Visual-Textual Alignment for Open Vocabulary Object Detection
This study introduces a novel Global and Local Visual-Textual Alignment method to improve open vocabulary object detection by effectively aligning image and text features. The approach enhances detector performance on unseen categories without requiring extensive pre-training data.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Vision-Language Models (VLMs) like CLIP are increasingly integrated into object detection frameworks, enabling open vocabulary object detection.
- Current methods often rely on large, uncurated image-text pairs for pre-training or use knowledge distillation, facing limitations like computational overhead and neglecting global information alignment.
- Existing approaches struggle with effective and efficient alignment between visual and textual features in the semantic space for open-set scenes.
Purpose of the Study:
- To propose a novel Global and Local Visual-Textual Alignment method for open vocabulary object detection.
- To address the limitations of current pre-training and knowledge distillation methods by integrating global and local feature alignment.
- To improve the ability of object detectors to perceive unseen objects by enhancing visual-textual feature alignment.
Main Methods:
- Developed a unified learning paradigm integrating global image-caption alignment and local region-prompt alignment.
- Global alignment uses contrastive learning between whole image and caption representations from CLIP's components.
- Local alignment focuses on matching region embeddings from CLIP's image encoder with textual token prompts from its text encoder, coupled with a prompt tuning strategy.
Main Results:
- The proposed method was implemented on Faster R-CNN and evaluated on OV-COCO and OV-LVIS benchmarks.
- Achieved clear improvements over existing methods on detecting novel categories.
- Demonstrated favorable performance compared to state-of-the-art approaches in open vocabulary object detection.
Conclusions:
- The Global and Local Visual-Textual Alignment method effectively enhances open vocabulary object detection by unifying global and local feature alignment strategies.
- The approach offers a parameter-efficient way to adapt VLMs like CLIP for downstream object detection tasks.
- This method provides a promising direction for improving the perception of unseen objects in computer vision systems.
Related Concept Videos
Design Example: Identifying the Locations of Monuments in the Field Using Global Positioning System Device
Design Example: Alignment of a Road Line Using GIS
Types of Global Positioning System Surveys
Structural Classification of Joints
A fibrous joint is where the adjacent bones are united by fibrous connective...
Field Application of Global Positioning System
Local Attraction