Reducing eagerness to please can help language models refuse harmful requests

Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

Computation and LanguageArtificial IntelligenceMachine Learning

Summary

Sometimes AI chatbots try too hard to agree with users, which can make them say yes to harmful requests. The authors studied how to reduce this eagerness—called sycophancy—by training language models differently. They found that lowering sycophancy doesn’t always make the models refuse harmful requests better, but it does help them resist pressure from users to say yes when refusing is important. This offers new insights about making AI safer and more reliable in tricky conversations.

What this means in practice

  • For ai safety engineers: Improve language model training to prevent eagerness to please causing unsafe responses under user pressure.
  • For chatbot developers: Design fine-tuning routines that help chatbots maintain refusals to harmful requests even when users push for compliance.

Authors

Xu Wang, Difan Zou, Xuansheng Wu

Abstract

Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit the acquisition of a target concept by supplying its associated activation during learning. Across three Qwen3.5 base models, we use sparse autoencoders (SAEs) to identify the top-ranked sycophancy feature from paired sycophantic and independent responses, then validate its behavioral influence through inference steering. We subsequently inject the selected feature during supervised fine-tuning on sycophantic targets. Positive injection reduces learned sycophancy after removal (by 62.0% relative to ordinary fine-tuning in 35B-A3B), whereas modest negative injection increases it. Unexpectedly, these reductions in sycophancy do not consistently improve direct refusal of harmful requests, motivating a narrower evaluation of the same harmful intents under user pressure. In this setting, ordinary fine-tuning on sycophantic responses substantially weakens refusal, while selected checkpoints trained with positive injection recover part of the loss, including approximately 95% in 35B-A3B. These findings show that persistent sycophancy reduction does not guarantee stronger direct refusal, while identifying recovery under user pressure as a distinct, conditional benefit of training intervention.