Related Experiment Video
Updated: Aug 27, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
7.8K
Guiding visual attention in deep convolutional neural networks based on human eye movements
Leonard Elia van Dyck1,2, Sebastian Jochen Denzler1, Walter Roland Gruber1,2
1Department of Psychology, University of Salzburg, Salzburg, Austria.
Frontiers in Neuroscience
|September 30, 2022
Summary
This study guided Deep Convolutional Neural Networks (DCNNs) using human eye-tracking data to alter visual attention during object recognition. Non-human-like attention models focused on different image areas, impacting face detection but not increasing overall human-likeness.
Area of Science:
- Computational Neuroscience
- Computer Vision
- Artificial Intelligence
Background:
- Deep Convolutional Neural Networks (DCNNs) show parallels with the human ventral visual pathway for object recognition.
- Recent deep learning advances may reduce this biological similarity, posing challenges for computational neuroscience.
- Previous work explored biologically inspired architectures; this study uses a data-driven approach.
Purpose of the Study:
- To investigate a purely data-driven method for guiding DCNN visual attention using human eye-tracking data.
- To assess if manipulating training data to mimic or contrast human visual focus impacts DCNN object recognition.
- To evaluate the effects on human-likeness and implications for face detection theories.
Main Methods:
- Human eye-tracking data was used to modify training images, guiding DCNNs' visual attention.
- Manipulation types included standard, human-like, and non-human-like attention focus.
- GradCAM saliency maps were compared against human eye-tracking data for validation.
Main Results:
- Guided focus manipulation effectively directed DCNN attention, particularly in the negative (non-human-like) direction.
- Non-human-like attention models focused on significantly different image regions compared to humans.
- Effects were category-specific, influenced by animacy and face presence, and emerged post-feedforward processing, notably impacting face detection.
Conclusions:
- While the data-driven attention guidance influenced DCNNs and face detection, it did not significantly increase human-likeness.
- The findings suggest potential applications for overt visual attention in DCNNs and offer insights into face detection mechanisms.
Related Concept Videos
Visual System
657
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
657
Vision
55.1K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.1K
Focusing of Light in the Eye
3.1K
Light rays enter the eye through the cornea, a transparent dome-shaped tissue that is the eye's outermost layer. The cornea bends or refracts, light rays traveling to the pupil. The shape of the cornea determines how much of the light is bent and whether the image will be focused correctly on the retina at the back of the eye. Once the light has passed through both refraction layers, it converges into a single focal point onto a small area. This is where photoreceptors start transforming...
3.1K
Depth Perception and Spatial Vision
847
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
847

