Related Experiment Video
Updated: Sep 10, 2025

High-resolution, High-speed, Three-dimensional Video Imaging with Digital Fringe Projection Techniques
Published on: December 3, 2013
Viewport prediction with cross modal multiscale transformer for 360° video streaming
Yangsheng Tian1, Yi Zhong2, Yi Han3
1School Of Information Engineering, Wuhan University of Technology, LuoShi Road, Wuhan, 430070, Hubei, China. 369487529@qq.com.
This study introduces a new model for predicting user views in 360° video streaming. The Cross Modal Multiscale Transformer (CMMST) improves prediction accuracy by considering user movement and visual importance.
Area of Science:
- Computer Science
- Multimedia Systems
- Artificial Intelligence
Background:
- Efficient 360° video streaming faces challenges with high bandwidth demands and unpredictable user viewports.
- Current methods often overlook inter-modal dependencies and personalized user preferences, leading to suboptimal prediction accuracy.
Purpose of the Study:
- To develop a novel viewport prediction model for 360° video streaming.
- To enhance prediction accuracy and efficiency by integrating diverse features and user-specific data.
Main Methods:
- A Cross Modal Multiscale Transformer (CMMST) model was developed.
- The model integrates user trajectory and video saliency features across multiple scales.
- Cross-modal attention mechanisms were employed to capture user preferences and viewing patterns.
Main Results:
- The CMMST model demonstrated superior performance compared to baseline methods.
- High prediction precision was maintained even with longer prediction intervals.
- The model effectively captures intricate user preferences and viewing behaviors.
Conclusions:
- The proposed CMMST model offers a promising solution for adaptive 360° video streaming.
- This approach can significantly improve user experience in virtual reality and other immersive platforms.
- The integration of multimodal features and attention mechanisms is key to accurate viewport prediction.
Related Concept Videos
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
Vector Transformation in Rotating Coordinate Systems
Cross Product
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
Uniform Depth Channel Flow
Depth Perception and Spatial Vision
Scaling

