Related Experiment Video
Updated: May 22, 2026

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
Window-to-window BEV representation learning for limited FoV cross-view geo-localization.
Lei Cheng1, Daikun Liu1, Lingquan Meng1
1School of Automation, Southeast University, Nanjing, 210096, Jiangsu, China.
Summary
This study introduces Window-to-Window Bird
Area of Science:
- Computer Vision
- Geospatial Artificial Intelligence
Background:
- Cross-view geo-localization faces challenges with significant perspective changes and limited ground-view field of view (FoV).
- Unknown ground image orientation and missing camera parameters create ambiguity in matching ground features to Bird's Eye View (BEV) representations.
Purpose of the Study:
- To develop a novel framework for learning Bird's Eye View (BEV) representations directly from ground features, addressing challenges of unknown orientation and limited FoV.
- To improve the accuracy and robustness of cross-view geo-localization systems.
Main Methods:
- Proposed a Window-to-Window BEV (W2W-BEV) representation learning framework.
- Implemented adaptive window-scale matching between BEV queries and ground references.
- Utilized depth-guided BEV initialization to project ground features into BEV space for enhanced matching.
Main Results:
- The W2W-BEV framework adaptively matches BEV queries to ground references at a window scale.
- BEV features are generated by attending to semantically relevant ground features within matched windows.
- Demonstrated significant superiority over state-of-the-art methods on benchmark datasets, particularly under challenging conditions.
Conclusions:
- W2W-BEV effectively bridges the cross-view domain gap by learning from limited FoV ground imagery.
- The proposed method enables learning aligned BEV features by borrowing information from ground views.
- Achieved superior performance in cross-view geo-localization with unknown orientation and limited FoV.
Related Concept Videos
Depth Perception and Spatial Vision
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
Spherical Coordinates
Spherical coordinate systems are preferred over Cartesian, polar, or cylindrical coordinates for systems with spherical symmetry. For example, to describe the surface of a sphere, Cartesian coordinates require all three coordinates. On the other hand, the spherical coordinate system requires only one parameter: the sphere's radius. As a result, the complicated mathematical calculations become simple. Spherical coordinates are used in science and engineering applications like electric and...
Relative Motion Analysis using Rotating Axes
Consider a component AB undergoing a linear motion. Along with a linear motion, point B also rotates around point A. To comprehend this complex movement, position vectors for both points A and B are established using a stationary reference frame.
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it instrumental in...
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it instrumental in...
Deconvolution
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Relative Motion Analysis using Rotating Axes-Problem Solving
Consider a crane whose telescopic boom rotates with an angular velocity of 0.04 rad/s and angular acceleration of 0.02 rad/s2. Along with the rotation, the boom also extends linearly with a uniform speed of 5 m/s. The extension of the boom is measured at point D, which is measured with respect to the fixed point C on the other end of the boom. For the given instant, the distance between points C and D is 60 meters.
Here, in order to determine the magnitude of velocity and acceleration for point...
Here, in order to determine the magnitude of velocity and acceleration for point...
Orthogonal Trajectories
Orthogonal trajectories describe the geometric relationship between two families of curves that intersect each other at right angles. One illustrative case involves a family of parabolas that open sideways along the x-axis. These curves share a common shape but differ by a scaling parameter, resulting in a set of curves that all pass through the origin and widen at different rates.Determining Orthogonal TrajectoriesTo identify the orthogonal trajectories for these parabolas, the first step...
