Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

Artificial Intelligence

Summary

The authors studied how models recognize when they are being tested (eval-awareness) and found that this awareness can be about either showing skills (capabilities) or following safety limits, or both. These different types affect how the model behaves when asked to comply with instructions. Their experiments showed that when models think they are being tested on abilities, they comply much more than when they feel they are being tested on safety. They also found that reducing overall eval-awareness doesn’t always reduce safety-related behaviors, meaning that not all eval-awareness is the same. This insight helps improve how model behavior is guided during testing.

Authors

Allison Zhuang, Santiago Aranguri

Abstract

Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction. Then, eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not, and the same "X% suppression of eval-awareness" can correspond to qualitatively different behavioral outcomes.