Transformer model improves long multivariate time series forecasting accuracy

SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

Machine Learning

Summary

Forecasting many related measurements over long times is hard because the data is complex and high-dimensional. The authors developed a transformer model called SETTer that uses special attention and masking techniques to better find important patterns in both time and different data channels. SETTer also includes simple explainable parts that show what patterns it focuses on. Their tests on multiple real datasets show it beats many current top methods most of the time.

What this means in practice

  • For energy grid operators: Use SETTer to predict long-term power usage patterns more accurately for operational planning and demand management.
  • For financial trading teams: Apply SETTer to forecast multiple financial indicators over extended periods to improve trading strategies and risk assessment.

Authors

Abraham Ezema, Chijioke Eze, Ferdinanda Ponci, Antonello Monti

Abstract

Long-term multivariate time series plays a significant role in many application areas such as power systems, trading, etc. However, their accurate prediction is quite difficult for conventional forecasting methods as they often exhibit high dimensionality and complex relationships. Recent works show that transformer-based approaches are quite effective for long-term forecasting thanks to their attention mechanism. However, in the presence of complex high-dimensional inputs, they show evidence of oversmoothing, limited capacity, and opacity. To this end, this paper introduces SETTer, a transformer-based model that addresses these challenges by incorporating novel techniques for decoupled self-attention and hybrid masking. The proposed techniques enable SETTer to effectively capture the dominant short- and long-term patterns across the temporal and channel dimensions. In addition, we enrich the model layers with simple explainable structures that indicate the discriminative pattern of SETTer. We show that with a single-layer transformer architecture, SETTer can effectively model long-term dependencies in the presence of varying data complexities. Extensive experiments on real-word benchmark datasets for long-term multivariate time series forecasting demonstrate that SETTer outperforms state-of-the-art models in 88% of the scenarios.