Related Experiment Videos
Scanpath Prediction in Panoramic Videos Via Expected Code Length Minimization
None:
Scanpath prediction in panoramic videos is a challenging task due to the spherical geometry and multimodality of the input, and the inherent uncertainty and diversity of the output. To give a complete treatment of these characteristics, we first present a simple criterion for scanpath prediction based on principles from lossy data compression. This criterion suggests minimizing the expected code length of quantized scanpaths, corresponding to fitting a discrete conditional probability model via maximum likelihood. We condition the probability model on two modalities: a viewport sequence as the deformation-reduced visual input and a set of relative past scanpaths projected onto respective viewports as the aligned path input. Furthermore, we parameterize it by a product of discretized Gaussian mixture models to capture the uncertainty and diversity of scanpaths from different humans. In doing so, the training of the probability model does not rely on the specification of "ground-truth" scanpaths for imitation learning. We also introduce a proportional-integral-derivative (PID) controller-based sampler to generate realistic human-like scanpaths from the learned probability model. Experimental results demonstrate that our method consistently produces better quantitative scanpath results in terms of prediction accuracy (by comparing to the assumed "ground-truths") and perceptual realism (through machine discrimination) over a wide range of prediction horizons. We additionally verify the perceptual realism improvement via a formal psychophysical experiment and the generalization improvement on several unseen panoramic video datasets.