Runtime system detects and fixes errors in multi-agent workflows

ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems

Artificial IntelligenceMachine Learning

Summary

Complex software workflows often use several specialized agents and tools working together, but problems during the tasks can stop everything from finishing correctly. The authors introduce ResonAct, a system that watches these workflows as they run, spotting problems by tracking continuous signals about progress and tool health. It then figures out the cause and automatically tries to fix errors without needing to change the original workflows. Testing shows this approach helps complete more tasks successfully with modest extra computing effort.

What this means in practice

  • For enterprise workflow teams: Monitor and automatically repair complex multi-agent workflows to improve task completion and reduce downtime during long-running business processes.
  • For cloud service operators: Deploy a control system that watches and fixes faults in orchestrated tool and agent interactions without modifying application code or orchestration logic.

Authors

Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta, Soujanya Soni, Sameep Mehta

Abstract

Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that prevent task completion. While existing observability frameworks provide traces and logs, diagnosis and remediation are largely performed after execution completes, limiting opportunities for recovery during runtime. We present ResonAct, a runtime self-healing framework that enables continuous monitoring, diagnosis, and remediation of multi-agent systems through streaming operational metrics. ResonAct ingests execution traces, agent interactions, and tool invocations into a streaming analytics layer that continuously derives task progress, context health, and tool reliability metrics. These metrics serve as runtime control signals for detecting anomalous execution patterns and localizing root causes using a structured failure model. Based on the diagnosed failure, ResonAct dynamically selects remediation policies and performs actions. The framework operates as an external control plane, enabling intervention without modifying application agents or orchestration logic. We evaluate ResonAct across enterprise workflow scenarios and AppWorld benchmarks. The results show that the streaming metric-based analysis identifies execution degradations and localizes faults. Furthermore, policy-driven remediation improves task completion rates by up to 10.00 percentage points, with detection precision ranging from 70.59% to 82.91%, recall from 63.09% to 100%, recovery rates from 10.48% to 46.67%, and runtime overhead ranging from $-0.25%$ to 14.12% across the evaluated configurations.