FlowState improves long task memory and lowers costs for AI agents

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Artificial Intelligence

Summary

Long tasks need AI agents to remember earlier parts of their work, but keeping all past information is expensive and squeezing it loses important details. The authors introduce FlowState, which stores and updates a special memory that keeps track of important facts and evidence separately. It lets the AI bring up old information only when needed and update decisions as new information arrives. This method improved success rates and cut down the token use by about 40% compared to remembering everything all at once.

What this means in practice

  • For ai agent developers: Build more efficient AI systems that handle complex tasks by selectively recalling and updating memory instead of storing entire histories.
  • For enterprise automation teams: Deploy AI workflows that save on processing costs while improving task success rates through smarter memory use.

Authors

Minghao Li, Bangyan Li, Zifan Wang, Yulong Li, Hu Xu, Gan Zhang, Jingtong Wu, Wenqiang Xu

Abstract

Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.