Neural network weights optimized to use smaller variable grammar codes

Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights

Machine Learning

Summary

Deep learning models use many numbers called weights, which can take up a lot of space. The authors show a way to tweak these weights so they can be described by shorter, repeating patterns called grammars, saving memory. They do this by grouping similar weight values together and training the model to work well with those simplified weights. This makes the weight representation smaller but with only a small drop in accuracy.

What this means in practice

  • For deep learning engineers: Optimize model weight storage by compressing weights into smaller variable-length grammar patterns to save memory with limited accuracy impact.
  • For embedded systems developers: Deploy compressed neural network weights in edge devices by using grammar-based encoding to reduce model size and resource use.

Authors

Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà

Abstract

We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.