Related Experiment Videos
PeNorm: Enhancing positional discriminability in transformers through normalization-based encoding optimization
Xilong Zhang1, Ruochen Liu1, Xuefeng Liang2
1School of Artificial Intelligence, Xidian University, No. 266 Xifeng Road, Xi'an, 710126, Shaanxi, China.
Abstract:
Positional encoding is a critical component in Transformer-based models, enabling them to capture sequential order in the absence of recurrent or convolutional structures. Despite its widespread use, practical implementations often exhibit limitations in distinguishing fine-grained positional information, particularly when sequences contain subtle or repetitive structural patterns. In this work, we first identify an empirical deficiency through qualitative and quantitative analyses: under commonly used positional encoding schemes such as sinusoidal and learnable absolute positional encoding, position vectors for nearby or structurally similar tokens can become insufficiently distinguishable, impairing the model's ability to differentiate positional contexts. To address this limitation, we propose PeNorm-a simple yet effective normalization-based optimization that enhances the discriminability of positional encodings by promoting uniformity and separation in the embedding space. We provide both qualitative and quantitative analyses demonstrating that PeNorm improves the distinctiveness of positional representations. Furthermore, we establish theoretical justification by showing that a three-layer Transformer Encoder, when equipped with appropriate positional encoding, can achieve perfect accuracy on the Parity-Odd language-a formal task requiring precise positional tracking and counting-thereby establishing a strong link between effective positional encoding and computational expressiveness. Empirical evaluations on formal language recognition tasks, machine translation benchmarks, and language modeling tasks consistently show performance improvements when using PeNorm, validating its efficacy in strengthening positional awareness. Our findings bridge a gap between theoretical capability and practical implementation, offering a principled and practical enhancement to improve the robustness of sequence modeling in Transformers.
Related Concept Videos
Transformers with Off-Nominal Turns Ratios
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Energy Losses in Transformers
There are four main reasons for energy losses in transformers.
The first cause can be the high resistance of the copper windings...
The Normal and Binormal Vectors
Maximizing the Directional Derivative
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...