From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised

Abstract