Sub-goal guided reflection improves long task success for language agents

SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning

Artificial Intelligence

Summary

Large language models struggle with long, complex tasks because they can lose track of important steps along the way, leading to mistakes. The authors propose a new approach that breaks down big goals into smaller sub-goals and guides the model to reflect on these intermediate steps. This method helps the model stay on track and complete complicated tasks more reliably, showing big improvements in several test environments. The technique also explores multiple plans to find better strategies, making language agents better at handling long sequences of actions.

What this means in practice

  • For interactive game developers: Improve AI agents in video games to complete complex multi-step in-game tasks more reliably using structured sub-goal planning.
  • For robotics teams: Enhance robot control systems by integrating sub-goal guided language planning to reduce errors in executing long sequences of actions.

Authors

Jaeho Jung, Sung Hoon Jung

Abstract

Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLACT, which reflects only on the end-goal at each step without explicitly considering intermediate sub goals. To address this problem, we propose SGG-ReflAct (Sub-Goal Guided Re flAct), a reasoning backbone that integrates sub-goals generated through a single path LLM planner into the reflection process. We further extend this framework to BeamSGG-ReflAct, which replaces the single-path planner with a beam search based LLM planner for structured plan exploration. We run experiments on ALF World, ScienceWorld, and Jericho with multiple LLM models. SGG-ReflAct out performs REFLACT in nearly all settings, achieving best success rate gains of 14.9 percentage points on ALFWorld and 8.0 percentage points on ScienceWorld with Llama-3.1-8B-Instruct. Our experimental analysis shows that SGG-ReflAct re duces hallucinated actions and achieves its largest gains on procedurally ordered tasks. Furthermore, experimental results with BeamSGG-ReflAct show that the backbone's effectiveness depends on plan quality: explicitly specifying the re quired operations recovers gains that plan searching alone cannot achieve. These results demonstrate that SGG-ReflAct offers a practical and highly effective rea soning backbone, enabling LLM agents to achieve reliable performance in com plex, long-horizon tasks through easy integration.