On the Threat Model of Weird Generalization and Emergent Misalignment
2026-08-24 • Computation and Language
Computation and Language
AI summaryⓘ
The authors studied a strange effect called weird generalization (WG), where fine-tuning a model on small, specific datasets can cause unexpected changes in its behavior. They found that WG depends more on what kind of data is used and the language it’s in than on how much data there is. WG happens more with data the model has seen during its original training and is also affected by the questions used to test it. The authors suggest WG is fragile and likely needs careful handling to avoid problems, rather than being a common issue in normal fine-tuning.
fine-tuningweird generalizationdomain-specific datasetmodel behaviordataset compositionpretrainingevaluation sensitivitylanguage modelsparametric knowledge
Authors
Miriam Wanner, Mark Dredze, William Walden
Abstract
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive this measurement is to the set of questions used. Experiments with three open-weight models on four datasets show that the degree of WG (1) depends heavily on dataset composition and language (more than on size); (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used. Collectively, these results indicate that WG is a product of quite fragile properties of both training and evaluation data. As such, we argue that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.