CacheDyG speeds up learning on dynamic graphs with less memory
CacheDyG: Decoupling Temporal Propagation for Efficient Dynamic Graph Learning
Machine Learning
Summary
Dynamic graphs represent things that change over time, like social networks or traffic patterns, and analyzing them takes a lot of computing power. The authors found that current methods waste time recalculating parts of the graph that don't change much between training steps. Their solution, CacheDyG, keeps a special memory bank of past computations and only updates small parts needed for learning. This makes training faster and uses fewer parameters while still making good predictions.
What this means in practice
- •For social media platform engineers: Accelerate training of models that predict user interactions evolving over time by reducing redundant calculations.
- •For network operations teams: Improve efficiency in learning from time-sensitive network traffic graphs by caching historical node representations.
Authors
PinHeng Zong, Ye Yuan
Abstract
Dynamic graphs are widely used to model time-evolving relational systems in real-world applications. Dynamic graph neural networks provide an effective framework for capturing both structural dependencies and temporal dynamics in such data. However, they typically intertwine temporal graph propagation with every optimization epoch and often maintain large trainable representations for each node-time pair. This design repeatedly recomputes largely unchanged historical structures, leading to substantial training and parameter overhead. To address this critical issue, we propose CacheDyG, a Cache-refine framework for efficient Dynamic Graph learning. Specifically, it decouples temporal propagation from routine parameter updates by constructing a time-ordered temporal dependency cache that stores graph-aware node-time representations in non-trainable buffers. During standard training epochs, CacheDyG reads from the cache and updates only a lightweight cache refiner, an adaptive residual gate, and the link predictor. Selective cache refresh further keeps cached representations aligned with the supervised objective while avoiding epoch-wise sparse propagation. Experiments on five dynamic graph benchmarks show that CacheDyG adopts substantially fewer trainable parameters and lower runtime to obtain more competitive predictive performance than baselines. These results demonstrate that cache-based decoupling provides an effective principle for scalable dynamic graph learning.