Language models learn to communicate efficiently using internal key value states
Draft-KV: Learning Useful Latent Communication Between Language Models
Computation and Language
Summary
When two language models work together, they can share information not just by speaking plain text but by exchanging hidden internal signals. The authors found that simply improving the receiver model’s accuracy doesn’t mean it used these signals properly. They created Draft-KV, which sends hidden key-value data generated during answer preparation, improving communication accuracy a lot with a small, specialized interface. This method lets a smaller model understand a much larger one’s knowledge without changing the models themselves. It even helps on new tasks and when models have access to different information.
What this means in practice
- •For ai engineers: Build communication layers between fixed large and small language models to improve overall performance on knowledge-intensive tasks.
- •For software developers: Create applications where smaller deployed language models access knowledge from larger models efficiently without retraining the main models.
Authors
Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang, Peng Zhong, Fengming Zhu, Xi Peng, Linqi Song, Jacky Keung, Jingyu Zhang
Abstract
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.