What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces

2026-08-24Computation and Language

Computation and Language
AI summary

The authors studied how changes made to a model’s internal activations affect its answers, but they noticed that results can depend on how the answers are labeled or encoded. They developed a new method to test interventions while changing answer labels, showing that models often follow the position of information in data (extraction index) rather than the meaning (semantic label). This behavior varies with different datasets and models. They conclude that evaluating model control needs careful testing because results can differ based on answer encoding.

activation steeringanswer encodingcontrastive activation additionextraction indexsemantic labelsmodel interventionNormBank datasetInference-Time Interventionbehavioral evaluationlanguage models
Authors
Zhiwei Gao, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki
Abstract
Activation steering is often evaluated under the answer encoding used to construct the direction. A reported gain may reflect the intended judgment or compatibility with answer identifiers seen during construction. We introduce Cross-Encoding Steering Evaluation, which freezes an intervention while re-encoding answers to the same held-out items. On NormBank, after A/B/C identifiers are reassigned, contrastive activation addition (CAA) induces larger target-versus-source score changes for the extraction indices than for the semantic labels under the new mapping. We call this extraction-index following. Varying identifier vocabulary (A/B/C, X/Y/Z, or 1/2/3) and row order shows that the effect tracks extraction index rather than row position. After matching direction norms across layers, extraction-index following emerges mainly at later depths. A low-rank output-sensitive component containing 15.4% of the direction's squared norm retains 96.3% of this effect. An Inference-Time Intervention (ITI)-style method also favors extraction-index over semantic-label following on NormBank in three models. In aggregate, MNLI favors extraction-index following, whereas Social Chemistry 101 (SC101) favors semantic-label following. Multiple-choice and open-ended evaluations can yield different behavioral conclusions. Thus, a steering gain under one answer encoding does not by itself identify what the intervention controls.