Single layer language model offers fast and memory efficient token prediction

Single-Layer MeMo as a Randomized Hamming-Kernel Classifier

Machine Learning

Summary

Predicting the next word in a sequence is a common task in language models. The authors study a simpler version of a model called MeMo that remembers how words follow each other using a simple matrix. They found this model works like a special kind of pattern matcher based on how different the positions in the input are. Their analysis explains how this model’s design affects prediction errors and shows it can be efficiently run on GPUs. Tests on a Wikipedia dataset show it balances prediction accuracy, speed, and memory use well compared to usual methods.

What this means in practice

  • For machine learning engineers: Implement a token prediction system that balances speed, memory, and accuracy on GPU-enabled hardware using MeMo’s single-layer approach.
  • For embedded system developers: Use a simplified language model for next-token prediction in memory- or compute-constrained devices where efficient retrieval matters.

Authors

Alessandro Straziota

Abstract

MeMo (Zanzotto et al., 2025) is a recent language-model architecture that stores associations between token contexts and next tokens in a correlation matrix memory. In this work, we study its single-layer form and show that its ideal retrieval rule is a multiclass classifier based on the positional Hamming kernel. The MeMo architecture represents both the sequence features and the output labels with Gaussian random codes. Its score is therefore a doubly randomized sketch of the ideal classifier. Under independent input and output codebooks, we bound the errors introduced by context sketching and output decoding, characterize their dependence on model and data parameters, and give a margin-based guarantee for recovering the ideal prediction. Controlled simulations support the trends predicted by the analysis. On a restricted WikiText-2 next-token task, we compare single-layer MeMo with classical baselines and show that it can offer a useful trade-off among predictive accuracy, memory, and throughput, particularly on a GPU, where its matrix operations can be parallelized.