Related Experiment Video
Updated: Jan 9, 2026

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective Diffusion.
This study introduces a new method for procedure planning in instructional videos, addressing uncertainty in visual observations and action selection. The approach enhances action prediction by synthesizing intermediate states and constraining action spaces for improved performance.
Area of Science:
- Computer Science
- Robotics
- Artificial Intelligence
Background:
- Procedure planning in instructional videos is complex due to limited visual data and vast action possibilities.
- Existing methods often implicitly handle uncertainty, leading to suboptimal performance.
Purpose of the Study:
- To develop an explicit solution for procedure planning that tackles uncertainty in visual observations and decision spaces.
- To improve the accuracy and efficiency of predicting action sequences for instructional videos.
Main Methods:
- Utilized image generation models and a prompt selection module within a diffusion model to synthesize diverse intermediate visual states.
- Introduced a task-selective diffusion model with a task-specific mask to constrain the action space.
- Enhanced visual representation using pre-trained vision-language models for action-aware, text-enriched multimodal embeddings.
Main Results:
- The proposed approach demonstrated superior performance on benchmark datasets compared to prior methods.
- The combination of synthesized states and a constrained action space significantly improved procedure planning accuracy.
- Action-aware multimodal embeddings enhanced task classification and subsequent action prediction.
Conclusions:
- The developed method effectively mitigates uncertainty in visual observations and decision spaces for procedure planning.
- This explicit approach offers a significant advancement in generating accurate and contextually relevant action sequences for instructional videos.
- The findings have implications for robotics, AI-driven instruction, and automated task execution.
More Related Videos
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018