Related Experiment Video
Updated: Apr 14, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.7K
Visual Object Recognition with 3D-Aware Features in KITTI Urban Scenes
J Javier Yebes1, Luis M Bergasa2, Miguel Ángel García-Garrido3
1Department of Electronics, University of Alcalá, Alcalá de Henares 28871, Spain. javier.yebes@depeca.uah.es.
Sensors (Basel, Switzerland)
|April 24, 2015
Summary
This study introduces 3D-aware features from stereo vision for detecting cars, pedestrians, and cyclists. The enhanced detection method improves performance on the KITTI benchmark, advancing autonomous driving perception.
Area of Science:
- Computer Vision
- Robotics
- Autonomous Systems
Background:
- Driver assistance systems and autonomous robotics require robust environment perception.
- Vision sensors offer cost-effective 3D scene understanding compared to LiDAR.
- Detecting road participants like cyclists and pedestrians in urban settings is crucial for navigation.
Purpose of the Study:
- To detect and estimate the orientation of cars, pedestrians, and cyclists using stereo color images.
- To develop 3D-aware features that capture both appearance and depth information.
- To extend the Deformable Part Model (DPM) for improved 3D object detection.
Main Methods:
- Computed 3D-aware features from stereo color images.
- Extended the part-based object detector (DPM) to utilize 2.5D data (color and disparity).
- Conducted a detailed analysis of the training pipeline and evaluated performance on the KITTI dataset.
Main Results:
- Achieved increased detection ratios for cars and cyclists compared to a baseline DPM.
- This work is the first to report results using stereo data for the KITTI object detection challenge.
- The proposed 3D-aware features enhance the understanding of road scenes.
Conclusions:
- The proposed method effectively utilizes stereo vision for 3D object detection in autonomous driving scenarios.
- The enhanced DPM with 3D-aware features shows competitive performance on challenging urban datasets.
- This research contributes to advancing the perception capabilities of autonomous systems.
More Related Videos
Related Concept Videos
Depth Perception and Spatial Vision
2.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.7K
Vision
61.7K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
61.7K

