Time series models get faster with compact adaptive context prompts
Instance-Adaptive Prompts as Context for Time-Series Foundation Models
Machine Learning
Summary
When predicting future values in time series data, using a long history helps but takes more computing power. The authors developed PaCTS, a method that creates a small set of adaptive tokens representing key information from the history instead of using all the data directly. These tokens capture both overall trends and short-term changes and can be used with existing frozen models to improve predictions. PaCTS works well across different datasets and models, making forecasting more accurate without needing much more computation.
What this means in practice
- •For data engineers: Speed up time series forecasting pipelines by replacing long input histories with compact adaptive prompts for faster inference without losing accuracy.
- •For energy grid operators: Improve demand forecasting with more accurate time series models that handle variable context length efficiently, aiding operational decisions.
Authors
Zehao Xiao, Shifeng Xie, Lei Zan, Jianfeng Zhang, Lujia Pan, Ievgen Redko, Malik Tiomoko, Keli Zhang
Abstract
Longer histories can improve time-series foundation models (TSFMs), but require substantially higher inference cost. We therefore ask whether contextual information can be provided more efficiently through a compact set of learned token embeddings. We introduce PaCTS, which generates a small set of instance-adaptive latent prompts in the form of continuous embedding tokens conditioned on the visible context. These prompts serve as compact context surrogates for frozen TSFMs. PaCTS constructs them from instance-specific global statistics and further refines them with segment-level temporal information, capturing both global characteristics and local temporal variations. The prompt module is jointly trained and deployed across heterogeneous time series with the frozen backbone. Extensive experiments demonstrate the effectiveness of prompts as context, consistently improving forecasting across context lengths and model architectures. With a shorter input context, PaCTS can outperform the same frozen backbone using double context while requiring substantially less inference computation. Compared with weight-space adaptation methods, PaCTS achieves stronger improvements and better out-of-distribution generalization.