Papers for

api service providers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Distillation defenses fail after language models undergo reinforcement learning

Distillation Defenses Easily Break After Reinforcement Learning

Abstract: Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.

Mon 28 SeptMachine LearningArtificial IntelligenceCryptography and Security
The gist
Bad actors can copy the reasoning skills of big language models by training their own models on the original model's responses, a process called distillation. The authors found that defenses designed to stop this copying seem to work only if attackers stop training right after. But if attackers keep training their model using reinforcement learning, these defenses are easily broken. This means existing protections might give a false sense of security. The authors suggest new defense ideas that work at a group level of model responses may be more effective.
Open → 2609.35699v1