Related Experiment Videos
Information-Theoretic Analysis of Positional Encoding Strategies in Vision Transformers: A Comparative Study of Four
Abstract:
Vision Transformers (ViTs) rely on positional encoding (PE) because self-attention has no native notion of token order or image-grid location, yet the information-theoretic properties of different PE strategies and their downstream consequences for model behaviour remain insufficiently characterised. We present a systematic comparison of four PE approaches-Learned, Sinusoidal, Rotary Position Embedding (RoPE), and a 1D-ALiBi-style linear-bias variant-together with a targeted 2D-ALiBi-style diagnostic intervention motivated by a raster-distance mismatch diagnosis. ViT-Base models are trained on three low-to-intermediate data-regime datasets (CIFAR-100, TinyImageNet, and ImageNet-100), using a primary cross-method matrix and a canonical paired protocol for the targeted 2D-vs-1D ALiBi-style comparison. We apply a two-track diagnostic suite: embedding-space analyses (per-dimension variance, entropy, PCA, and probes) for additive PE, and attention-space intrinsic analyses (bias-tensor rank and entropy, slope schedule, and RoPE short-wavelength band counting) for attention-space PE, combined with attention-position mutual information (MI), noise ablation, and PE removal. The results reveal four qualitatively distinct encoding regimes: Learned PE is "quiet and ubiquitous," Sinusoidal PE "loud and structured," RoPE attains the highest accuracy with a graceful mid-depth MI decay, and the ALiBi-style variant exhibits a persistent attention-space bias. Direct attention-space analysis shows that the 1D-ALiBi-style bias tensor has full intrinsic rank. PE removal reveals a ${\sim }10\times$ dependency spectrum-from Sinusoidal collapse to near-chance to Learned PE retaining most of its accuracy-indicating extensive implicit positional learning. Linear probes show that Sinusoidal PE gives $0\%$ held-out column accuracy-a protocol-level generalisation failure for the modulo-column label rather than absence of column information-while Learned PE partially recovers 2D structure. Identifying the ALiBi-style raster-scan distance $|i{-}j|$ as mismatched to the 2D geometry of image patch grids, we introduce 2D-ALiBi-style, which replaces it with the 2D Euclidean patch-grid distance. On the canonical CIFAR-100 paired cohort ($n{=}12$ seeds), fixed-slope 2D-ALiBi-style yields a statistically reliable $+0.45$ pp accuracy improvement ($p{=}0.013$), is directionally positive on TinyImageNet, and shows no statistically reliable difference on ImageNet-100. It also lowers CLS-excluded patch-only attention-position MI. The mean-magnitude-matched 2D arm reverses this MI effect and reaches $70.28\pm 0.35\%$, the highest accuracy among the three canonical ALiBi-style arms; a secondary within-2D paired contrast ($+1.94$ pp, $12/12$ positive) establishes bias magnitude as a substantial lever without supporting a magnitude-rather-than-geometry interpretation. We therefore frame the contribution as a controlled characterisation plus a modest, statistically supported intervention, not a state-of-the-art PE method. Overall, PE choice affects accuracy, robustness, and attention-space information flow beyond what accuracy alone captures.
Related Concept Videos
Color Vision
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Gestalt Principles of Perception
Three-Winding Transformers
In the per-unit equivalent circuit of a grounded Y-Y three-phase...