Language models balance fast output and traceable text with new sampling method

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

Cryptography and SecurityMachine Learning

Summary

Large language models can create text quickly and let people know that the text came from them. But doing both at the same time is hard. The authors designed a new way to pick words using a process based on random timing, which helps keep the text fast and easy to check. Their method also makes sure the text carries a special signature (a watermark) without slowing things down. They tested their idea and found it works well for both speed and watermark quality.

What this means in practice

Authors

Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz Gündüz

Abstract

Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have shown that combining these two goals is highly nontrivial and can be potentially impossible. In this work, we develop a novel multi-draft speculative sampling algorithm based on Poisson processes that improves the frontier of this fundamental trade-off. The proposed algorithm has strong sampling efficiency on its own and, more interestingly, is naturally watermarkable: we can embed an unbiased watermark without degrading speculative acceptance. Moreover, our algorithm is based on an exact list-coupling-without-communication scheme, which yields a drafter invariance property that benefits both sampling and watermarking. It is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency, and we experimentally verify its strong performance in both aspects.