Formally verified data links assertions to behavior in hardware designs

EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations

Machine Learning

Summary

Often, computer tools generate checks called assertions for hardware designs based on descriptions in human language. However, these checks may accidentally rely on specific details of one version of a design instead of its overall behavior. The authors created a large, carefully checked dataset containing multiple different hardware designs that all do the same thing, along with matching assertions verified to capture the actual behavior. This lets people test whether a generated check truly reflects the intended function rather than incidental design choices. They also showed how to evaluate a popular AI model on this dataset to measure how reliable its generated checks are across different versions.

What this means in practice

Authors

FNU Aditi

Abstract

Large language models are increasingly used to generate SystemVerilog Assertions from natural-language specifica- tions and register-transfer-level designs. Existing datasets and benchmarks support important goals such as large- scale training, formal evaluation, specification-to-assertion generation, and mutation-based testing. A complemen- tary need is to study whether a generated assertion cap- tures externally observable behavior or depends on inci- dental details of one RTL implementation. We present EquivSVA, a formally verified dataset organized around behavior families. Each family contains four structurally distinct RTL implementations of the same externally ob- servable behavior, shared interface-level gold properties, three controlled mutants, and formal-validation evidence. EquivSVA contains 120 behavior families across 12 cat- egories, 480 reference RTL implementations, 914 gold properties, and 360 mutants. Every final family passes a fixed 17-job validation suite covering RTL equivalence, gold-property proofs, property reachability, mutant dis- tinguishability, and gold-property checks on mutants. We also provide fixed family-safe train, development, and test splits. As a small demonstration of the analyses en- abled by the dataset, we evaluate the publicly released, Apache-2.0-licensed Qwen2.5-Coder-7B-Instruct model on the held-out test split. Of 293 interface-only generated properties, 93 are formally sound, and the number of sound properties varies across equivalent implementations for 14 of 24 test families. These results illustrate how behavior-family organization can support controlled stud- ies of assertion-generation robustness without requiring changes in intended functionality. The dataset, generators, validation scripts, and case-study artifacts are publicly released at https://github.com/aditigupta96/EquivSVA.