Viewport prediction with cross modal multiscale transformer for 360° video streaming

Yangsheng Tian1, Yi Zhong2, Yi Han3

  • 1School Of Information Engineering, Wuhan University of Technology, LuoShi Road, Wuhan, 430070, Hubei, China. 369487529@qq.com.

Scientific Reports
|August 19, 2025
PubMed
Summary

This study introduces a new model for predicting user views in 360° video streaming. The Cross Modal Multiscale Transformer (CMMST) improves prediction accuracy by considering user movement and visual importance.

Related Concept Videos

Relative Motion Analysis using Rotating Axes01:25

Relative Motion Analysis using Rotating Axes

Consider a component AB undergoing a linear motion. Along with a linear motion, point B also rotates around point A. To comprehend this complex movement, position vectors for both points A and B are established using a stationary reference frame.
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
531
Vector Transformation in Rotating Coordinate Systems01:16

Vector Transformation in Rotating Coordinate Systems

Consider a vector rotating about an axis with an angular velocity, such that its tip sweeps a circular path.
1.8K
Cross Product01:25

Cross Product

The cross product is a fundamental concept in vector algebra that is a vector operation on two different vectors to obtain a third vector. Unlike the scalar product, the cross product results in a vector quantity perpendicular to both the original vectors.
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
330
Uniform Depth Channel Flow01:27

Uniform Depth Channel Flow

Uniform depth channel flow keeps fluid depth consistent along channels such as irrigation canals. In natural channels, such as rivers, approximate uniform flow is often assumed. This condition occurs when the channel’s bottom slope matches the energy slope, balancing potential energy lost from gravity with head loss due to shear stress. This balance prevents depth changes along the channel length, resulting in a steady, uniform flow.Uniform flow in open channels with a constant cross-section...
144
Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
900
Scaling01:26

Scaling

In designing and analyzing filters, resonant circuits, or circuit analysis at large, working with standard element values like 1 ohm, 1 henry, or 1 farad can be convenient before scaling these values to more realistic figures. This approach is widely utilized by not employing realistic element values in numerous examples and problems; it simplifies mastering circuit analysis through convenient component values. The complexity of calculations is thereby reduced, with the understanding that...
315