Papers for

enterprise ml security teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SpecGuard detects hidden backdoors in large language models instantly

SpecGuard: Inference-Time Backdoor Detection For Free

Abstract: Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger appears. While backdoors can be audited before deployment, runtime monitoring remains important for models that are frequently updated. The challenge is that LLM serving is latency-sensitive: existing inference-time detectors either rely on assumptions about the trigger form, which can fail on stealthy attacks, or require extra model computation, such as input perturbations or an additional generation pass. We introduce SpecGuard, an inference-time backdoor detector that repurposes speculative decoding at zero added model-computation cost. Speculative decoding speeds up inference by using a small draft model to propose tokens and a target model to verify them. We observe that this verification process already exposes a useful signal: when a backdoor is triggered, the target model shifts toward the attacker's behavior, while a clean draft model does not predict this shift, causing the draft-token acceptance rate to change. We formalize when this signal appears and show that an attacker who suppresses it must also weaken the backdoor. Across diverse backdoor types and model families, SpecGuard reliably detects triggered behavior, including stealthy cases where input-level filters are blind, while avoiding the extra generation cost of existing runtime detectors. Speculative decoding therefore doubles as a free, always-on signal for detecting backdoored LLM behavior.

Thu 10 SeptCryptography and SecurityComputation and Language
The gist
Large language models can have hidden "backdoors" that behave normally but change behavior when triggered by secret inputs. The authors created SpecGuard, a method that detects these backdoors during normal use without extra slowdowns. It works by comparing predictions from a smaller fast model with a bigger model; differences reveal malicious triggers. This method catches many types of backdoors quickly and cheaply, helping keep AI safer while running.
Open 2609.11799v1