Behave enables self improving agents for hardware design verification

BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification

Artificial IntelligenceHardware ArchitectureMachine Learning

Summary

Hardware designs can work correctly even if they operate at different speeds, which makes it hard to check if a design is right by looking at every step. The authors created BEHAVE, a tool that focuses on how hardware behaves rather than timing details, allowing designs to be verified more flexibly. BEHAVE trains an agent to improve hardware designs by testing and learning from different implementations and gradually solving harder tasks. This approach helped make hardware design agents better at verifying designs without needing exact time-step matches.

What this means in practice

  • For hardware design engineers: Create and verify hardware designs with flexible timing to explore better performance and resource trade-offs.
  • For ai system builders: Develop self-improving agents that learn from correctness feedback for complex hardware verification tasks.

Authors

Yuheng Wu, Berk Gokmen, Sujeeth Jinesh, Lauren McLane, Aarav Wattal, Qi Yang Huang, Zhaozhuo Xu, Thierry Tambe

Abstract

Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent iteratively develops a register-transfer-level (RTL) design and a behavior model as the design's verification reference. Our evaluator, BEHAVE-Sim, checks both artifacts separately against a hidden golden behavior model using input stimuli generated by random sampling and solver-guided search. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, the agent continually searches for high-level implementations relevant to its capability gaps, constructs and checks specification-behavior pairs, and trains on the expanded task pool. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. Starting from 60 seed tasks and acquiring 100 new tasks, self-improvement raises Qwen3.8-27B's RTL pass@1 on BEHAVE-Eval from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.