Reciprocal guidance improves token generation speed for ai models

Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier

Computation and Language

Summary

Generating text with AI models involves two main steps: drafting what to say next and then verifying if it fits well. The authors noticed that these two steps can actually help predict each other, so by using that insight, they created a system called Reciprocal Guidance. This system dynamically adjusts how much work to do on drafting and verifying depending on how busy the computer is. Their tests show it can speed up the text generation by nearly twice under different workload conditions.

What this means in practice

  • For ai service operators: Improve the efficiency of AI text generation services by dynamically balancing compute resources for drafting and verification stages under varying user loads.
  • For cloud infrastructure teams: Optimize allocation of compute resources in cloud systems that provide AI-based language generation to increase throughput and reduce latency.

Authors

Linye Wei, Shutian Zheng, Haoyu Zeng, Meng Li

Abstract

Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared backbone, while enabling longer acceptance lengths. However, the Pareto frontier between aggregate and per-request throughput remains underexplored. At low concurrency, sequential draft-verify execution requires two model forward passes per round, limiting the effective tokens per forward (TPF). By contrast, at high concurrency, longer drafts incur increasingly expensive computation, forcing individual requests to operate under constrained speculation budgets and preventing full exploitation of the full-backbone drafter. Our key observation indicates that drafting and verification exhibit reciprocal predictability. Draft logits can anticipate likely verification mismatches, while recent verification outcomes predict future drafting utility and suitable block sizes. Building on this observation, we introduce Reciprocal Guidance (RecGuide), a runtime draft-verify orchestration framework that adapts speculative decoding to varying serving loads. RecGuide exploits spare compute capacity through verification-overlapped drafting at low concurrency, while dynamically allocating request-specific draft block sizes as the workload becomes increasingly compute-intensive. Experiments across a wide range of concurrency levels demonstrate consistent throughput improvements over vanilla self-speculation, achieving up to $1.8\times$ speedup.