Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

709
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
709
Observational Learning01:12

Observational Learning

207
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
207
Gestalt Principles of Perception01:21

Gestalt Principles of Perception

340
Gestalt principles provide a framework for understanding how humans perceive objects as unified wholes within their context. These principles are essential in explaining the cognitive processes that make sense of complex visual stimuli by organizing them into coherent groups. One fundamental principle is proximity, which posits that objects located close to each other are perceived as a collective group. For instance, when dots are positioned near one another, the visual system interprets them...
340
Visual Agnosia01:12

Visual Agnosia

234
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
234
Purposive Learning01:22

Purposive Learning

139
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
139
Perceptual Constancy01:12

Perceptual Constancy

437
Perceptual constancy is the ability to recognize that objects remain consistent and unchanged even when their appearance varies due to changes in sensory input. There are four main types of perceptual constancy: size constancy, shape constancy, color constancy, and brightness constancy.
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
437

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective.

IEEE transactions on pattern analysis and machine intelligence·2026
Same author

Holistic Invariant Retracing for Distortion-Resilient Multi-Modal Learning in Spatial Transcriptomics.

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026
Same author

Demonstration of efficient predictive surrogates for large-scale quantum processors.

Nature communications·2026
Same author

GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification.

IEEE transactions on medical imaging·2026
Same author

A DeepSeek-powered AI system for automated chest radiograph interpretation in clinical practice.

Nature communications·2026
Same author

NoisePO: Efficient Semantic Noise Generation and Ranking for Diffusion-Based Text-to-Image Synthesis.

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

Related Experiment Video

Updated: Jul 16, 2025

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
07:12

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss

Published on: April 11, 2025

458

Learning Visual Affordance Grounding From Demonstration Videos.

Hongchen Luo, Wei Zhai, Jing Zhang

    IEEE Transactions on Neural Networks and Learning Systems
    |September 11, 2023
    PubMed
    Summary

    This study introduces a hand-aided affordance grounding network (HAG-Net) to improve human-object interaction segmentation. HAG-Net uses hand cues from videos to accurately identify interaction regions, outperforming existing methods.

    More Related Videos

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K
    Visualization Method for Proprioceptive Drift on a 2D Plane Using Support Vector Machine
    07:05

    Visualization Method for Proprioceptive Drift on a 2D Plane Using Support Vector Machine

    Published on: October 27, 2016

    9.3K

    Related Experiment Videos

    Last Updated: Jul 16, 2025

    Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
    07:12

    Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss

    Published on: April 11, 2025

    458
    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K
    Visualization Method for Proprioceptive Drift on a 2D Plane Using Support Vector Machine
    07:05

    Visualization Method for Proprioceptive Drift on a 2D Plane Using Support Vector Machine

    Published on: October 27, 2016

    9.3K

    Area of Science:

    • Computer Vision
    • Robotics
    • Human-Computer Interaction

    Background:

    • Visual affordance grounding identifies interaction regions between people and objects for applications like robot grasping.
    • Current methods struggle with multiple interaction possibilities on objects and multiple human actions within the same region.

    Purpose of the Study:

    • To propose a novel Hand-Aided Affordance Grounding Network (HAG-Net) to address limitations in current visual affordance grounding methods.
    • To leverage hand position and action cues from demonstration videos to refine interaction region localization.

    Main Methods:

    • HAG-Net employs a dual-branch architecture processing video and object data separately.
    • The video branch uses hand-aided attention and Long Short-Term Memory (LSTM) for action feature aggregation.
    • The object branch incorporates a Semantic Enhancement Module (SEM) and distillation loss for knowledge transfer.

    Main Results:

    • Evaluations on two datasets demonstrate HAG-Net achieves state-of-the-art performance in visual affordance grounding.
    • The method effectively disambiguates multiple interaction possibilities on objects and refines localization using hand cues.

    Conclusions:

    • HAG-Net successfully improves visual affordance grounding by integrating hand-centric information from demonstration videos.
    • The proposed network offers a more robust solution for segmenting human-object interaction regions in complex scenarios.