Related Experiment Videos
BiasFormer: Structured Posterior-Guided Transformer for Temporal Action Segmentation
Dongyue Zhou1, Zhihui Shi1, Haixia Wang1
1School of Computer and Artificial Intelligence, Beijing Technology and Business University, Beijing 100048, China.
Sensors (Basel, Switzerland)
|August 13, 2026
Summary
BiasFormer improves temporal action segmentation by integrating structural guidance and constraints. This unified framework enhances action boundary detection and temporal coherence in video analysis.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Transformer architectures excel at temporal action segmentation due to long-range dependency modeling.
- Current methods struggle with action durations and transitions, leading to fragmented and ambiguous segmentations.
Purpose of the Study:
- To introduce BiasFormer, a novel framework addressing limitations in transformer-based temporal action segmentation.
- To enhance temporal coherence and accuracy in action boundary detection.
Main Methods:
- Coupling feature-level structural guidance with decoder-level structural constraints.
- Utilizing frame-level confidence biases from structured Markov posteriors for feature reweighting.
- Implementing a unified segmental Markov head for explicit duration and transition modeling.
Main Results:
- BiasFormer demonstrates competitive performance across multiple datasets (Breakfast, GTEA, EgoProceL, EPIC-KITCHENS).
- Achieved improvements over existing methods like FACT, particularly in segment-level metrics.
- Validated the effectiveness of integrating posterior-guided feature refinement with structured decoding.
Conclusions:
- The proposed BiasFormer framework effectively improves segment-level temporal coherence in action segmentation.
- Coupling structural guidance and constraints offers a robust approach for refining temporal action segmentation models.