Related Experiment Video
Updated: Sep 17, 2026

End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
Cross-Modality Causal-Aware Hierarchical Representation Learning for Domain Generalized Object Detection
Abstract:
Domain Generalized Object Detection (DGOD) addresses the critical challenge of detecting objects across diverse unseen visual domains. Recent advances in vision-language models (VLMs) have shown promising zero-shot generalization capabilities that benefit DGOD. However, existing VLM-based DGOD methods primarily leverage VLMs for data augmentation, overlooking the rich generalizable knowledge it contains. Besides, VLM-based methods developed for other domain generalization (DG) tasks suffer from modality gap and representation bias due to their correlation-driven cross-modal interaction paradigms, which severely limit the performance to DGOD. To bridge this research gap and advance the causal-driven VLM-based DG, we develop a Causal-aware Hierarchical Representation Graph that reformulates the VLM-based DG problem into a hierarchical causal representation learning framework. Our framework incorporates a Generalizable Knowledge Transfer module to refine and inherit transferable scene-object features from the VLM's visual encoder, and a Causal Prototype Learning module that employs general text embeddings as causal interventions to guide the construction of a mediator visual causal prototype space, which inherits the generalization and category relational representation ability from text without representation bias. Furthermore, we introduce a Prototypical Cross-attention Classifier that eliminates modality gap by integrating object features with learned causal prototypes for text-free classification, which also directs visual features to approach the mediator causal space, enabling the learning of causal visual features that are invariant to domain-specific confounders. Our causal-driven framework transfers causal invariance from text to visual modality and provides a vision-friendly perspective for leveraging VLMs to solve vision-centric tasks. Extensive experiments on five benchmarks including Diverse Weather, Corruption, Real-to-Artistic, Cross Camera and Sim-to-Real demonstrate that our method achieves superior generalization performance.
Related Concept Videos
Associative Learning
Classical conditioning, also known...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...