Related Experiment Video
Updated: May 26, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Discovering thematic objects in image collections and videos
Junsong Yuan1, Gangqiang Zhao, Yun Fu
1School of Electrical and Electronics Engineering, Nanyang Technological University, Singapore. jsyuan@ntu.edu.sg
Summary
This study introduces a novel bottom-up method to automatically discover key thematic objects in images and videos. The approach efficiently identifies representative visual content despite variations, aiding search and summarization.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Image Analysis
Background:
- Identifying thematic objects in visual data is crucial for applications like object search, tagging, and video understanding.
- Existing methods face challenges due to a lack of prior knowledge about object characteristics and significant appearance variations.
Purpose of the Study:
- To develop a robust method for discovering representative thematic objects in image and video collections.
- To overcome limitations of top-down generative models by proposing a bottom-up approach.
Main Methods:
- A novel bottom-up strategy is proposed, involving gradual pruning of uncommon local visual primitives to recover thematic objects.
- A multilayer candidate pruning procedure is implemented to accelerate image data mining.
- The method is designed to handle objects of various sizes and tolerate significant appearance variations.
Main Results:
- The proposed solution effectively locates thematic objects across diverse datasets.
- The method demonstrates efficiency and robustness in identifying key visual content.
- Experimental results validate the effectiveness of the approach compared to existing methods.
Conclusions:
- The developed bottom-up approach offers an efficient and effective solution for thematic object discovery in visual data.
- This method enhances capabilities in object search, tagging, and video analysis.
- The technique shows promise for handling complex visual data with inherent variations.