Simple counts rival learned memory for predicting future links
Do Temporal Link Predictors Need Learned Memory? A Smoothed-Count Baseline with a Handful of Parameters
Machine LearningArtificial Intelligence
Summary
Predicting future connections between things that interact over time usually relies on complex methods that learn patterns from past interactions. This paper shows that just counting how often certain repeated interaction patterns occur, combined with some simple rules, can predict these future links just as well or better than many advanced methods. The authors created a model that uses these counts with only a handful of adjustable settings and found it works very well across a variety of datasets. This suggests that complicated learned memories might not always be necessary for good temporal link prediction.
What this means in practice
- •For social network engineers: Improve friend or contact recommendations by implementing stable, low-parameter methods for predicting new links without expensive node embedding training.
- •For communication platform developers: Use simple statistical counts to predict future interactions and optimize message routing or feature suggestions with minimal computational overhead.
Authors
Lisi Qarkaxhija, Ingo Scholtes
Abstract
Many temporal link predictors summarize past interactions through learned node representations. We examine whether simple counts of recurring interaction patterns can provide competitive predictions without learning these representations. We propose a temporal link predictor based on statistical language modelling. It pools transition and co-occurrence counts across sources to predict links that a source has never formed. We smooth sparse estimates using destination frequencies or Kneser-Ney continuation counts. A shared log-linear rule combines these estimates with popularity, source history, and recency, without node embeddings. In our main evaluation, the model achieves the highest MRR among the compared methods on 7 out of 16 datasets from TGB and TGB-Seq. It also outperforms EdgeBank and Base3 on all 16 datasets and the heuristic family on 14. These gains extend to datasets designed to limit repeated edges. With only 9--13 learned parameters, our model provides a simple and competitive baseline for evaluating future neural temporal link predictors.