Large reasoning models avoid false alarms by ignoring prompt tricks

DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models

Artificial Intelligence

Summary

Large reasoning models often say no too much when asked sensitive questions, partly because they learn easy but misleading clues from training prompts. The authors found two main shortcut clues: certain prompt formats and specific words that wrongly trigger refusals. Their method called DeShortcut-Align helps the models ignore these shallow cues by measuring which words influence refusals, creating new examples without harmful keywords, and training the models to be consistent even if prompt styles change. This makes models safer and less likely to reject innocent queries unnecessarily while keeping their reasoning skills strong.

What this means in practice

  • For ai safety engineers: Improve the reliability of refusal mechanisms in large models to reduce false rejections while maintaining safety guardrails.
  • For language model developers: Develop robust fine-tuning and reinforcement learning pipelines that minimize reliance on superficial prompt cues for better generalization.
  • For content moderation teams: Deploy safer large models that avoid excessive denial of benign content triggered by simple keyword or formatting traps.$Commercial implications: Enables improved moderation tools with fewer false positives for social media platforms and online services.

Authors

Qirui Liu, Yichen Sun, Yan Wang, Zhixuan Chu, Linbo Jiang, Jianan Lin, Kui Ren

Abstract

Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.