Stratified prompting reduces unwanted bias while keeping good model traits
Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
Artificial Intelligence
Summary
When teaching language models, sometimes they learn bad habits along with good ones. The authors found that asking the model to show the bad habit during training and then telling it not to during use still lets bad habits slip out. Their new approach, stratified inoculation prompting, uses a small set of clean examples in different situations to help the model keep good behaviors without the bad ones. This method cuts down unwanted behaviors more and keeps more good behaviors than before, even in tricky situations.
What this means in practice
- •For machine learning engineers: Improve fine-tuning workflows to reduce harmful or undesired model outputs while keeping intended capabilities intact.
- •For ai safety teams: Deploy models less likely to produce harmful advice or unwanted behaviors even under unexpected prompts.
Authors
Kajetan Dymkiewicz, Tim Farrelly, Adam Práda, Ishaan Panigrahi, Srishti Gureja, Helen Yannakoudakis, Robert Mullins, Victor Gillioz, Daniel Tan, Maxime Riché
Abstract
Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompts. IP can also hinder learning of the desired behaviour. We address these limitations in settings where both behaviours co-occur in most training examples, so filtering out examples with undesired behaviour leaves only a small clean subset. We introduce stratified inoculation prompting (SIP). SIP leverages a small clean subset to demonstrate that desired behaviour should persist without the undesired one across different contexts. SIP oversamples these clean examples under diverse non-eliciting prompts while inoculating the rest. SIP substantially reduces expression of undesired behaviour while preserving more of the desired behaviour than IP. These gains persist even when we extend IP to oversample the same clean subset at the same rate as SIP. Moreover, SIP yields lower emergent misalignment rates in all harmful-advice setups we tested. SIP can be further extended to limit the undesired behaviour even under prompts that explicitly request it. We introduce backdoor dilution, which weakens expression under the inoculation prompt, and password-locked inoculation, which concentrates elicitation on a designated password. Taken together, our findings show that changing the training contexts for a small clean subset can significantly improve selective generalisation.