Unified 16-bit format boosts neural network error protection efficiency

UniCASE: A Unified 16-bit Floating-Point Format with Criticality-Aware Selective ECC for Efficient DNN Protection

Hardware ArchitectureCryptography and Security

Summary

Soft errors can cause mistakes in how deep learning models work, which harms their accuracy. The authors propose UniCASE, a new way to store numbers that both saves space and protects against these errors more efficiently. Their method looks at which parts of the number are most important and protects them better, reducing the overhead of error correction. Tests show UniCASE keeps accuracy close to standard formats while cutting error correction costs and improving reliability.

What this means in practice

  • For deep learning engineers: Implement more reliable neural network models by integrating efficient error protection that reduces overhead without losing accuracy.
  • For hardware designers: Design floating-point units with selective error correction that balance protection strength and computation cost for AI accelerators.

Authors

Amna Hassan, Semeen Rehman

Abstract

Soft errors are an increasing reliability concern for Deep Neural Network execution because they can corrupt parameters, leading to accuracy degradation. While conventional ECC offers strong fault protection, it incurs additional parity storage and computational overhead. Embedded-parity formats reduce storage cost by reusing the least-significant bits, but they do not optimize protection while reducing computational overhead. We propose UniCASE, a unified 16-bit floating-point (FP) format that jointly optimizes data representation and error protection for reliable DNN execution. It identifies stable blocks across FP64, FP32, FP16, and BFloat16 that can be mapped into a unified representation. Based on bit-level criticality analysis, UniCASE uses selective ECC that assigns distinct levels of protection to different data bits according to their resilience against soft errors. Experimental results show that UniCASE reduces encoder/decoder cost by up to 30%, preserves model accuracy within 1% of the FP32 baseline, and provides significantly stronger soft error resilience than existing embedded-parity methods.