Multimodal time series model separately handles internal and external data
QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities
Machine Learning
Summary
Forecasting future events from time series data can be challenging because some data types affect the system from within (endogenous), while others influence it from the outside (exogenous). The authors propose QiYao-M, a model that treats these two kinds of data differently to improve prediction accuracy. It learns how internal data changes over time and uses a special retrieval method to adapt quickly to external data, even with limited training examples. Tests show this approach performs well across various forecasting tasks.
What this means in practice
- •For energy grid operators: Improve electricity demand and supply forecasts by separately modeling internal usage patterns and external factors like weather.
- •For financial analysts: Enhance stock and commodity price predictions by better distinguishing between intrinsic market trends and external economic indicators.
Authors
Hanyin Cheng, Linfeng Wang, Zhengbo Qu, Yang Shu, Zhongwen Rao, Meng Wang, Yijie Li, Xin Jiang, Bin Yang, Chenjuan Guo
Abstract
Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two types of modalities separately. For endogenous modalities, to capture how they evolve along with the underlying temporal dynamics, we introduce an Endo-Multimodal Predictor and Endo-Multimodal Supervision to explicitly learn their evolution from history to the future. For exogenous modalities, to generalize across domains and across various modality types and numbers under the scarcity of exo-multimodal pretraining data, we propose an Exo-Multimodal Retrieval Enhancer that enables rapid downstream adaptation without updating the TSFM parameters. We further introduce Endo-Modality Proxy Training to train this retrieval module without exogenous multimodal pretraining data. Extensive experiments across unimodal and multimodal benchmarks demonstrate strong forecasting performance in scenarios both with and without exogenous modalities.