Odin speeds up encrypted llama 3 inference on nh100 gpu

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Cryptography and Security

Summary

Cloud services that run large language models usually require sending your text so the provider can see it, which risks your privacy. The authors developed Odin, a system that lets these models run on encrypted input without revealing your data. They improved how data is packed and processed inside encryption, making inference much faster and using less memory. Odin runs Llama 3 8B on a powerful NVIDIA GPU much quicker than previous methods, with the same privacy protections.

What this means in practice

  • For cloud service operators: Enable large language model inference on encrypted user inputs to protect privacy without sacrificing performance.$Commercial implications: This technique enables cloud providers to offer privacy-preserving AI services, attracting clients concerned about data confidentiality.
  • For security software engineers: Incorporate optimized FHE layouts and computations to accelerate privacy-preserving neural network inference on GPUs.

Authors

Yuhang Fan, Yusi Chen, Kanyu Ye, Zhuoran Ji

Abstract

Cloud LLM services typically require users to send prompts to a model provider, creating a privacy risk. Fully homomorphic encryption (FHE) lets a server perform inference without decrypting the input, but representing data as ciphertexts adds storage and computational overhead. In CKKS-based LLM inference, the packing scheme maps logical tensors to ciphertexts and slots. It therefore determines the ciphertext count and the homomorphic cost of linear layers, and it constrains how data pass between linear layers, attention, and nonlinear computation. As models and sequences grow, inefficient layouts accumulate encoding, compute, and layout-conversion overhead. We present Odin, an FHE inference system that co-designs ciphertext packing and model execution for Llama. Starting from a THOR-style baseline whose bottleneck is weight encoding, Odin uses a feature-major cross-layer layout to unify residual connections and layer interfaces, and builds transient intra-operator layouts for linear projections and attention. This reduces redundant plaintext encoding of weights in wide projections. Within attention, QK^T produces scores that Softmax can consume directly, and PV consumes the resulting probabilities, avoiding intermediate repacking. For nonlinear ops, we use minimax polynomial approximation with input-range control and joint error allocation guided by model quality, reducing polynomial degree and multiplicative depth. To our knowledge, Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3. With Llama-3-8B weights and a 128-token input, Odin evaluates all 32 Transformer layers on a single NVIDIA H100 80 GB GPU. Server-side end-to-end FHE evaluation takes 366.4 s and 58.9 GiB peak device memory. Under the same model, input, CKKS parameters, and hardware, THOR takes 1651.9 s, a 4.51x speedup.