DuplexCadence speeds up real-time speech with less memory use

DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines

Artificial Intelligence

Summary

Turning spoken words into computer-generated speech in real time is tricky because everything must happen fast and on a strict schedule. The authors found that the slowest part causes delays and wastes memory by handling many small tasks inefficiently. They created DuplexCadence, which better matches the model’s internal timing to the computer’s operation, reducing wasted memory and speeding up processing. This approach lets speech models keep up more reliably with live conversation while using fewer resources.

What this means in practice

Authors

Haixiao Gao, Yimin Zheng, Linyou Xiao, Zeke Xie

Abstract

Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches $2.85\times$ the stock runtime's speed at $38.8\%$ lower peak memory. On the live duplex path, mean SPEAK time falls from $14\%$ over the one-second cadence to $2\%$ under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-