Fiona speeds up encrypted neural network inference with smarter weight use

Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights

Machine Learning

Summary

Running AI models on encrypted data keeps information private but is very slow. The authors show that changing how certain weights are simplified can drastically speed up this process without making the AI much less accurate. They created an optimizer called Fiona that decides which parts of the model to simplify and which to keep precise, balancing speed and accuracy. Tests on three popular AI models showed Fiona could double inference speed while keeping accuracy loss under 1%.

What this means in practice

  • For cloud ai service providers: Deploy faster privacy-preserving AI models that process encrypted user data efficiently using Fiona's selective weight simplification technique.$Commercial implications: Fiona enables commercially viable encrypted AI inference services by significantly reducing computation time while maintaining model accuracy.
  • For security-focused ai engineers: Improve the performance of secure neural network inference by integrating selective ternarization and hybrid weight strategies based on Fiona's approach.

Authors

Yiteng Peng, Zhibo Liu, Dongwei Xiao, Shuai Wang

Abstract

Fully homomorphic encryption (FHE) enables neural network inference directly on encrypted inputs, but it remains orders of magnitude slower than plaintext in- ference. Applying the server's plaintext weights to encrypted activations involves plaintext-ciphertext multiplications (PMult) and accounts for more than half of inference time in recent systems. Ternary quantization can replace these multipli- cations with additions and subtractions, but the savings rarely materialize under packed execution. A single PMult applies a weight group fixed by the packing layout and can be avoided only when all its weights share the same ternary value. Ternarizing all groups, however, largely degrades accuracy. We present FIONA, an offline optimizer that selectively ternarizes weights within a given packing layout based on the estimated effect of ternary conversion on the model's performance. FIONA encourages a shared ternary value within each weight group and retains full-precision weights for sensitive groups, so ternar- ized and full-precision paths coexist within a layer. It then compiles these hybrid operators exactly, applying common scaling factors once to accumulated inputs and reusing sums across outputs. Weight ternarization can also narrow the input ranges of downstream polynomials. FIONA fits lower-degree replacements under a cumulative accuracy budget, reducing multiplicative depth and bootstrapping. On VGG11, ViT, and BERT, FIONA reduces PMult operations by 53.4-79.5% and accelerates end-to-end encrypted inference by 2.38x, 1.68x, and 1.84x, re- spectively, with less than 1% accuracy loss across all three models.