The Fragility of Jailbreak Robustness Across Operational States
2026-08-31 • Cryptography and Security
Cryptography and SecurityComputation and Language
AI summaryⓘ
The authors found that the success of tricks (or 'jailbreaks') used to bypass safety rules in language models can change a lot just by tweaking ordinary system prompts, even if those prompts aren't supposed to affect safety. They tested this on different models and attacks and saw big changes in how often the tricks worked, sometimes more than doubling the success rate. They also discovered that these changes are linked to differences in the model's internal behavior related to refusal responses. Their work suggests that testing with only one standard setup doesn't give a complete picture of how resistant models are to these tricks.
jailbreak attacksattack success rate (ASR)system promptoperational statealigned modelsmodel robustnesshidden representationsrefusal responselanguage modelsmodel evaluation
Authors
Yuna Park, Hwang Youn Kim, Yujin Kim, Won Woo Ro, Suhyun Kim, Jae-In Hwang
Abstract
Existing jailbreak evaluations typically characterize robustness using a single attack success rate (ASR) measured in a default configuration (the vanilla state). However, user-LLM interactions can induce diverse operational states beyond the vanilla state. In this work, we find that jailbreak robustness is highly fragile to operational-state variation: even when the attack remains fixed, changing only an ordinary system prompt not designed to affect safety can dramatically alter attack success rates. We systematically investigate this phenomenon across seven aligned models and three representative jailbreak attacks, observing substantial differences in ASR between vanilla and non-vanilla operational states. In one case, ASR increases by up to 56 percentage points (2% to 58%) solely due to a change in operational state. Remarkably, these increases occur even for attacks originally designed and optimized under vanilla-state evaluation. We further show that state-dependent robustness variation is systematically associated with differences in hidden representations along a refusal-related axis, and that projections onto this axis strongly predict jailbreak outcomes. Our results show that a single vanilla-state evaluation may not fully characterize jailbreak robustness, motivating evaluations that also examine how robustness changes across non-vanilla operational states.