Related Experiment Video
Updated: Jan 29, 2026

Topographical Estimation of Visual Population Receptive Fields by fMRI
Published on: February 3, 2015
Pose-Perceptive Convolution: Learning Geometry-Adaptive Receptive Fields for Robust 6D Pose Estimation
Yi Lai1, Yaqing Song1, Qixian Zhang2
1College of Information, Mechanical and Electrical Engineering, Shanghai Normal University, Shanghai 201418, China.
None:
6D object pose estimation is crucial for applications such as robotic manipulation and augmented reality, yet it remains highly challenging when dealing with objects of significantly different aspect ratios or the drastic appearance variations of a single object caused by pose changes. Most existing methods focus on designing more complex backend fusion modules, while largely overlooking a fundamental problem at the feature extraction frontend: the geometric mismatch between the fixed, square receptive fields of standard convolutions and the varied projected morphologies of objects. This mismatch, along with noise in fused features and ambiguity in regression, limits the performance ceiling of current methods. To this end, this paper proposes a novel Pose-Perceptive Convolution (PPC) and constructs a new Pose-Perceptive Fusion Network (PPF-Net). Its core component, the Pose-Perceptive Convolution, fundamentally resolves the aforementioned geometric mismatch by dynamically adapting the shape and sampling density of its receptive field. Experiments on four benchmarks show that PPF-Net improves the VSD score by 19.4% over FFB6D on MP6D, and achieves 96.7% ADD-S on YCB-Video, approaching state-of-the-art accuracy. Crucially, these gains are realized with minimal computational overhead, avoiding the heavy latency of backend-intensive approaches. This validates that frontend feature extraction is an efficient strategy for robust 6D pose estimation.
Related Concept Videos
Coordination Number and Geometry
Predicting Molecular Geometry
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
Geometry of Hyperbolas
Subliminal Perception
Factors Affecting Perception
An illustrative example of a perceptual set is the scenario where an airline pilot told...

