Related Experiment Video
Updated: May 26, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
8.9K
SAVE: Self-Attention on Visual Embedding for Zero-Shot Generic Object Counting
Ahmed Zgaren1,2, Wassim Bouachir2, Nizar Bouguila1
1Concordia Institute for Information Systems Engineering (CIISE), Concordia University, Montréal, QC H3G 1M8, Canada.
Journal of Imaging
|February 25, 2025
Summary
This study introduces an automated zero-shot counting method that surpasses existing zero-shot and few-shot techniques. The novel approach enhances visual object counting accuracy for diverse applications.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Generic Visual Object Counting aims to identify and quantify objects within images.
- Zero-shot counting enables object counting for arbitrary classes without prior examples, contrasting with few-shot methods that require exemplars.
- Existing methods often require exemplars or lack the automation needed for rapid processing.
Purpose of the Study:
- To propose a fully automated zero-shot counting method that outperforms current zero-shot and few-shot approaches.
- To enhance the accuracy and efficiency of visual object counting across various domains.
Main Methods:
- Exploiting feature maps from a pre-trained detection-based backbone.
- Introducing a Visual Embedding Module to generate semantic embeddings with object contextual information.
- Utilizing a Self-Attention Matching Module to create an encoded representation for the head counter.
Main Results:
- Achieved state-of-the-art performance in zero-shot counting on the FSC147 dataset.
- Obtained the best Mean Absolute Error (MAE) of 8.89 and Root Mean Square Error (RMSE) of 35.83.
- Demonstrated competitive results compared to few-shot methods.
Conclusions:
- The proposed method offers a significant advancement in automated zero-shot visual object counting.
- The approach shows promise for applications in tree counting, wildlife monitoring, and medical image analysis (e.g., blood cell counting).
- This work pushes the boundaries of visual object counting, enabling more efficient and accurate automated solutions.
Related Concept Videos
Depth Perception and Spatial Vision
510
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
510
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K

