ZonoGPT verifies large transformers with depth-independent method

ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models

Machine LearningSoftware Engineering

Summary

Verifying that large AI models behave safely and fairly is difficult, especially for big transformer networks like GPT. The authors introduce ZonoGPT, a new technique that checks these models without using more memory as the models get deeper. They achieve this by cleverly reducing complexity and keeping important relationships in the data during verification. Their approach works on popular, large GPT models and proves that the models satisfy certain properties on many examples.

What this means in practice

  • For ai safety engineers: Verify robustness and safety properties of deployed GPT-based systems using a scalable verification method.
  • For large ai model developers: Improve trust in large transformer models by formally checking correctness and fairness before release.

Authors

Hai Duong, Thanh Le, ThanhVu Nguyen

Abstract

Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, safety, and fairness, neural network verification techniques prove required properties and provide auditable guarantees before deployment. However, prior work remains limited to small or restricted Transformers, and maintaining precision across deep models remains challenging. In this work, we introduce ZonoGPT, an abstract domain for verifying large transformers that maintains a space complexity independent of network depth. ZonoGPT uses a structured zonotope and a generator reduction mechanism to efficiently preserve correlations. To maintain precision, it introduces block-specific fused transformations for Attention and LayerNorm that retain feature relations, along with an affine transform for GELU that preserves generator relations. These mechanisms enable \tool{} to be the first approach to verify standard architectures, scaling to official HuggingFace models up to GPT-2 Medium (24 blocks, 300M+ parameters) and successfully verifying 1,339 instances across text and vision tasks.