Lifted training improves classifier accuracy without changing network design

ProtoSeam: Lifting Classifier Training with Latent Gaussian Mixture Models

Machine Learning

Summary

Training machine learning classifiers can sometimes be tricky because the way we train them affects how well they perform later. The authors propose a new training method that inserts special reference points called prototypes inside the network during training. These prototypes help the earlier part of the network produce outputs closer to the right class, making the whole system more accurate. When using the model later, these prototypes are removed, so the network works exactly like before but with better accuracy.

What this means in practice

  • For machine learning engineers: Improve accuracy of existing image classification models without altering their runtime architecture by incorporating lifted training with prototypes.
  • For computer vision developers: Deploy more accurate vision models on standard datasets like CIFAR and TinyImageNet using lifted training to boost performance without changing model inference.

Authors

Robert Lampel, Timon Klein, Sebastian Sager

Abstract

We propose a lifted reformulation of supervised classification that improves the final accuracy of standard classifiers without changing the architecture at inference time. A network $N=N_2\circ N_1$ is split at a single semantic interface and one learnable prototype per class is inserted there. Training combines a quadratic consensus penalty that pulls $N_1(x)$ toward the prototype of its class with a classification loss of $N_2$ evaluated on samples drawn around the prototypes, whereat no gradient crosses the interface. At inference the prototypes are discarded and the unmodified network $N_2\circ N_1$ is used. Across CIFAR-10, CIFAR-100, and TinyImageNet with ResNet and vision transformer backbones, lifted training improves test accuracy by up to five percentage points over variants without lifting under a shared tuning protocol. Moreover, we provide theoretical justification of those results.