Related Experiment Video
Updated: Jan 30, 2026

Novel Object Recognition Test for the Investigation of Learning and Memory in Mice
Published on: August 30, 2017
Object-guided contrastive language-image pre-training for zero-shot target recognition
1Southwest China Institute of Electronic Technology, Chengdu, 610036, Sichuan, China. zch.512@163.com.
Abstract:
Target recognition is critical for security systems, but traditional Visual-Language Models (VLMs) like CLIP suffer from limited training data semantics, poor background suppression, and inflexible multi-resolution features. To address these, we propose Object-Guide CLIP (OG-CLIP), integrating three core enhancements: Knowledge graph-driven data augmentation: A 5000-category military knowledge graph and 1M image-text pairs via multi-source acquisition and knowledge-infused prompts. Target-centered ROI module: Fuses SAM 2-generated masks with ViT features to focus on discriminative regions and suppress background noise. Adaptive MRL: Resolves traditional MRL's rigid granularity via 128D-1024D continuous features, dynamic dimension weighting, and cross-granularity semantic alignment. Experiments on 99 target categories (military aircraft, warships, civilian targets) show OG-CLIP achieves 84.28% mean Accuracy (mAcc), 11.36 percentage points higher than baseline CLIP. Ablation confirms contributions of each component, and OG-CLIP excels in complex scenarios. The proposed framework offers a scalable and adaptable vision-language modeling approach for military recognition, with future work focusing on dataset expansion and model lightweight optimization.
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
pre-mRNA Processing
Once about 20-40 ribonucleotides have been joined together by RNA polymerase, a group of enzymes adds a “cap” to the 5’ end of the growing transcript. In this process, a 5’ phosphate is replaced by modified guanosine that has a methyl group attached to it (7-Methyl...
Velocity of an Object

