DuplexCadence speeds up real-time speech with less memory use
DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines
Artificial Intelligence
Summary
Turning spoken words into computer-generated speech in real time is tricky because everything must happen fast and on a strict schedule. The authors found that the slowest part causes delays and wastes memory by handling many small tasks inefficiently. They created DuplexCadence, which better matches the model’s internal timing to the computer’s operation, reducing wasted memory and speeding up processing. This approach lets speech models keep up more reliably with live conversation while using fewer resources.
What this means in practice
- •For voice assistant developers: Improve voice assistants by enabling faster, more memory-efficient real-time speech interaction that keeps up reliably with users.
- •For real-time translation software teams: Enable translation systems to produce speech with less delay and memory usage during simultaneous listening and speaking.
Authors
Haixiao Gao, Yimin Zheng, Linyou Xiao, Zeke Xie
Abstract
Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches $2.85\times$ the stock runtime's speed at $38.8\%$ lower peak memory. On the live duplex path, mean SPEAK time falls from $14\%$ over the one-second cadence to $2\%$ under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-