Llm agents learn to forget bad action sequences safely

Trajectory Unlearning on LLM-based Agents

Artificial Intelligence

Summary

Large language models (LLMs) being used as autonomous agents can sometimes repeat unwanted behaviors, not just say harmful things. This paper introduces the idea of removing specific action sequences, or trajectories, that an agent should forget, which is different from erasing knowledge. The authors propose a new method called GiRPO that helps agents forget these bad action paths while still performing well on desired tasks. They tested their approach on household and online shopping tasks and showed it works better than previous methods.

What this means in practice

  • For robotics developers: Remove harmful or unsafe robot behaviors by selectively forgetting bad action sequences without affecting overall task success.
  • For ecommerce platform engineers: Prevent autonomous shopping agents from repeating undesired browsing or purchasing behaviors by unlearning specific action paths.
  • For chatbot service providers: Improve conversational assistants by deleting sequences of actions or responses that lead to poor user experiences.$Commercial implications: Enables chatbot companies to sell enhanced agent customization by removing unwanted behaviors from AI assistants.

Authors

Yingdan Shi, Ren Wang

Abstract

Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an agent knows, an agent should not reproduce undesired behaviors through its action trajectories. In this work, we introduce trajectory-level unlearning, a new problem formulation that targets the removal of specific action trajectories in long-horizon agentic tasks, rather than factual knowledge. We identify two fundamental challenges that distinguish trajectory unlearning from knowledge unlearning: (1) our unlearning target is what the agent \emph{does}, not what it \emph{says}; and (2) trajectories are sequentially dependent action sequences that cannot be decomposed into isolated prompt-response pairs without losing inter-step structure. To address these challenges, we propose Group-injected Relative Policy Optimization (GiRPO), which injects forget trajectories into the policy rollout group with penalized rewards and isolates the normalization statistics, yielding a stable and bounded unlearning signal that does not corrupt gradient updates for normal task trajectories. We construct trajectory unlearning benchmarks from two application scenarios, household tasks (ALFWorld) and online shopping (WebShop), and design three complementary metrics for evaluating forgetting quality and model utility. Experiments on ALFWorld and WebShop demonstrate that GiRPO effectively unlearns target trajectories while preserving task success rates, outperforming existing knowledge-unlearning baselines on both forgetting quality and task utility.