Related Experiment Video
Updated: Aug 6, 2026

06:54
Photorealistic Learned Landscapes for Augmented Reality
Published on: June 27, 2025
FreeScene: Generative-enhanced visual modeling for generalized zero-shot image retrieval with scene sketches
Ran Zuo1, Zhengming Zhang2, Chenxu Ji3
1Communication University of China, Beijing, 100024, China.
Summary
This study introduces FreeScene, a novel framework for zero-shot sketch-based image retrieval (ZS-SBIR) at the scene level. FreeScene enhances retrieval accuracy by integrating multimodal large language models and diffusion models for improved semantic and structural understanding.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Existing zero-shot sketch-based image retrieval (ZS-SBIR) methods primarily focus on instance-level retrieval and struggle with complex scene compositions.
- Fine-grained scene-level (FG-SL) methods face challenges in zero-shot settings due to intermingled seen and unseen object categories.
- Current approaches fail to effectively address the dual challenges of fine-grained scene-level matching and generalized zero-shot retrieval in complex scenes.
Purpose of the Study:
- To propose a novel framework, FreeScene, for scene-level zero-shot sketch-based image retrieval (ZS-SBIR).
- To address the limitations of existing methods in handling fine-grained scene-level matching and generalized zero-shot retrieval with mixed object categories.
- To enhance visual representations and zero-shot generalization by leveraging multimodal large language models (MLLMs) and diffusion models.
Main Methods:
- Developed FreeScene, a framework integrating semantic and structural cues from scene sketches and images.
- Utilized Multimodal Large Language Models (MLLMs) for distilling high-level scene semantics via textual embeddings.
- Implemented a semantic-centroid confidence strategy to align textual embeddings with visual features, injecting external semantic priors.
- Introduced a hybrid CNN-ViT backbone with an adaptive attention mechanism for multi-scale structure encoding.
- Employed a generative feature correlation module with diffusion-based denoising to optimize representations.
Main Results:
- FreeScene significantly outperforms state-of-the-art methods in scene-level ZS-SBIR.
- The framework effectively captures both semantic and structural information for improved retrieval.
- Demonstrated enhanced zero-shot generalization capabilities in complex scene retrieval tasks.
- Established a new benchmark for zero-shot learning in scene-level sketch-based image retrieval.
Conclusions:
- FreeScene provides a robust solution for the challenging task of scene-level zero-shot sketch-based image retrieval.
- The integration of MLLMs and diffusion models proves effective in enhancing semantic understanding and visual representations.
- The proposed methods successfully address fine-grained matching and generalized zero-shot retrieval in complex scene compositions.
- FreeScene sets a new standard for future research in scene-level ZS-SBIR and zero-shot learning applications.
Related Concept Videos
Shape and Texture of Coarse Aggregate
Aggregate shape is classified based on the relative sharpness or roundness of the edges and corners. This classification includes categories like rounded, angular, elongated, and flaky, each with specific characteristics. Rounded aggregates, fully shaped by attrition, are typical of river or seashore gravel, while angular aggregates, such as crushed rock, have well-defined edges. Aggregates that are elongated and flaky are less desirable, as they can reduce the workability and strength of...
Fischer Projections
Learning to draw Fischer projections of molecules and understanding their relevance plays a crucial role in the visual depiction of organic molecules. A Fischer projection is a two-dimensional projection on a planar surface to simplify the three-dimensional wedge–dash representation of molecules. This is especially helpful in the case of molecules with multiple chiral centers that can be difficult to draw. Here, all the bonds of interest are represented as horizontal or vertical lines. While...
