Neural network training verified cheaply without trusting trainers

Training Witnesses: Trusting the Training without Trusting the Trainer

Machine LearningCryptography and Security

Summary

Machine learning results depend on trusting the person who trained the model, which can lead to mistakes or unfair comparisons. The authors propose a way to check if training was done correctly without redoing the entire process. Their method uses quick behavioral checks and occasional tests to confirm the model's training and data use. This makes it easier and cheaper to verify models and helps create trusted benchmarks for progress.

What this means in practice

  • For machine learning engineers: Verify that large neural network training runs used the correct data and processes without costly retraining.
  • For machine learning platform operators: Maintain leaderboards of validated training runs to enable consistent baselines and track progress in model development.
  • For cloud service providers: Offer training certification services that audibly verify neural network training runs to customers.$Commercial implications: Certification of training correctness can be offered as a service to AI developers and enterprises requiring trusted model training.

Authors

Houjun Liu, Pratyusha Sharma

Abstract

Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.