Data aware rotary encoding improves token order understanding in transformers
RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings
Computation and LanguageMachine Learning
Summary
Transformers need a way to know the order of words or tokens because they process them without built-in order. The existing common method, RoPE, struggles with long-distance positions beyond the training context, making some predictions less accurate. The paper introduces DaRoPE, which keeps the strengths of RoPE for close tokens but learns positions differently for tokens far apart, based on the data context. This approach works better for sequences like music, DNA, and neural signals, and matches or beats RoPE in language tasks too.
What this means in practice
- •For natural language engineers: Improve transformer models’ understanding and generation of longer text sequences with better positional cues.
- •For bioinformatics teams: Enhance analysis of genomic sequences by better capturing long-range positional relationships in DNA data.
Authors
Jarod Lévy, Mathurin Videau, Jad Yehya, Jean-Rémi King, Stéphane d'Ascoli, Thomas Moreau
Abstract
Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing alternatives have been evaluated under different settings, leaving the literature fragmented and without a clear replacement. We bring structure to this landscape by examining a specific weakness of RoPE: its slow frequency bands, whose wavelengths exceed the training context and expose models to unseen angles during extrapolation. We therefore introduce Data aware RoPE (DaRoPE), which preserves standard RoPE on the fast bands but replaces absolute position on the slow bands with bounded coordinates learned from contextual representations. Therefore, the slow-band geometry depends on the data rather than only on positional distance. We compare representative encodings under matched conditions across synthetic tasks, symbolic music, genomics, neural signals, and language models spanning 124M to 50B parameters. Across these experiments, DaRoPE leads on non-text benchmarks, mitigates recency bias, while remaining best or on par in language modeling and length extrapolation. Moreover, the learned coordinates also make the mechanism interpretable, revealing how attention layers leverage contextual information beyond token distance. Together, these results support DaRoPE as the best overall default among the evaluated methods, when there is no domain-specific reasons to prefer another.