Atlas reduces language model changes to keep old answers intact
Learn Here, Move Less Elsewhere: Input-Conditioned Plasticity from Retained-Domain Activation Atlases
Machine LearningArtificial Intelligence
Summary
When language models are trained for new tasks, their old answers can unintentionally change, which can be a problem. The authors introduce a method called ATLAS that helps models learn new tasks without altering how they respond to previous ones. It uses a clever system of reference points and directional rules to guide updates, so the model keeps its original knowledge while learning. Tests showed ATLAS keeps answers more stable and accurate across different tasks and models.
What this means in practice
- •For software engineers: Implement task-specific updates to language models that preserve existing capabilities while adding new skills.
- •For ai application integrators: Maintain consistent AI system behavior during frequent domain-specific model customizations to reduce errors in deployed services.
- •For automated coding tool developers: Improve coding assistance models by enabling specialized training without overwriting stable previously learned code generation abilities.$Commercial implications: Allows creation of commercially viable coding assistants that adapt to new programming tasks while maintaining reliable older functionality.
Authors
Jiangtao Lin, Bangyang Wei, Yihang Ding, Siyi Liu, Yuhan Dong
Abstract
Task-specific fine-tuning can rewrite a language model's answers beyond the training task, complicating updates that must preserve existing behavior. We introduce ATLAS, which turns retained-domain representations into an input-dependent rule for task adaptation. An activation atlas supplies local reference centers and directional filters to a shared low-rank residual. Target supervision learns the residual, while retained geometry shapes its action throughout training and inference. On Qwen3-8B, ATLAS achieves lower mean retained-output Kullback-Leibler (KL) divergence than all seven published baselines at shared coding-performance requirements, with consistent advantages across multiple training seeds. Structural comparisons identify the contributions of retained reference states and directional conditioning, and answer-level analyses show fewer rewritten mathematical answers and more stable commonsense choices. Experiments spanning five backbones and two retained domains further demonstrate coding gains with reduced retained-output movement. With compact storage and modest decoding overhead, ATLAS provides a practical mechanism for acquiring specialized skills while maintaining continuity in existing responses.