Related Experiment Video
Updated: May 24, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
7.6K
Group Visual Relation Detection
Summary
This paper introduces Group Visual Relation Detection (GVRD) to identify relations involving groups in images. The proposed Simultaneous Group Relation Prediction (SGRP) method effectively detects these group visual relations (GVRs).
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Existing visual relation detection focuses on individual objects, overlooking group interactions.
- Groups are prevalent in image semantics, necessitating advanced detection methods.
- Group Visual Relation Detection (GVRD) extends traditional methods by including group subjects and/or objects.
Purpose of the Study:
- To introduce a novel task, Group Visual Relation Detection (GVRD).
- To propose a new method, Simultaneous Group Relation Prediction (SGRP), for addressing GVRD.
- To create and release a new dataset, COCO-GVR, for GVRD research.
Main Methods:
- The Simultaneous Group Relation Prediction (SGRP) method is proposed.
- SGRP comprises Entity Construction (EC), Feature Extraction (FE), and Group Relation Prediction (GRP) modules.
- The EC module generates instances and group/phrase candidates; FE extracts multi-modal features; GRP predicts groups and predicates simultaneously.
Main Results:
- The novel COCO-GVR dataset, containing 9,570 images and 31,855 GVRs, was created.
- Extensive experiments were conducted on the COCO-GVR dataset.
- The SGRP method demonstrated superior performance compared to baseline methods.
Conclusions:
- GVRD is a valuable extension to visual relation detection.
- The SGRP method effectively addresses the GVRD task.
- The COCO-GVR dataset facilitates future research in group visual relation detection.
More Related Videos
Related Concept Videos
Visual Agnosia
173
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
173
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K
Visual System
475
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
475
Gestalt Principles of Perception
269
Gestalt principles provide a framework for understanding how humans perceive objects as unified wholes within their context. These principles are essential in explaining the cognitive processes that make sense of complex visual stimuli by organizing them into coherent groups. One fundamental principle is proximity, which posits that objects located close to each other are perceived as a collective group. For instance, when dots are positioned near one another, the visual system interprets them...
269
Depth Perception and Spatial Vision
508
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
508
Relative Motion Analysis using Rotating Axes-Problem Solving
382
Consider a crane whose telescopic boom rotates with an angular velocity of 0.04 rad/s and angular acceleration of 0.02 rad/s2. Along with the rotation, the boom also extends linearly with a uniform speed of 5 m/s. The extension of the boom is measured at point D, which is measured with respect to the fixed point C on the other end of the boom. For the given instant, the distance between points C and D is 60 meters.
Here, in order to determine the magnitude of velocity and acceleration for point...
Here, in order to determine the magnitude of velocity and acceleration for point...
382

