Transformer models can handle long sequences with one size fits all
Universality and Generalization of Causal Transformers Across Context Lengths
Machine Learning
Summary
Modern transformer models often need different setups for different sequence lengths, but this work shows a single transformer can effectively handle sequences of varying lengths uniformly. The researchers studied how continuous and stable token sequences can be approximated by one fixed transformer, extending to infinitely long sequences modeled as continuous curves. They also provide a mathematical bound on how well these transformers generalize to new data without depending on maximum sequence length. Experiments on real-world time-series data support their mathematical assumptions, showing differences between physical signals and text embeddings.
What this means in practice
- •For natural language processing teams: Design transformer models that handle variable-length text inputs without reconfiguring or retraining for each sequence length.
- •For time series analysis teams: Use causal transformers to model long continuous signal data efficiently with uniform performance guarantees across different sampling resolutions.
Authors
Takashi Furuya, Maarten V. de Hoop, Gabriel Peyré
Abstract
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by $α$-Hölder sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a $β$-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error $O((\log\log N/\log N)^{β/(d+2)})$ from $N$ iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with $d$ the token dimension and no maximum-length factor. Finally, experiments on physical time series support the Hölder-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.