Beyond Stepwise Modeling: Toward a Unified Contextual Reasoning Framework for Hyperspectral Video Object Tracking