Papers for

workflow engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Ai agents face hidden failures calling tools in workflows

When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary

Abstract: AI agents increasingly execute long-running workflows that externalize effects through independently supplied tools. Under retries, speculative execution, concurrency, and partial failures, the resulting external state may be inconsistent with the workflow's intended resolution: required effects may be missing or duplicated, aborted effects may survive, and committed effects may depend on provisional state that is later withdrawn. Advanced transaction models address related failures, but assume that lower-level operations expose the semantics they depend on: whether an effect occurred, whether it can be compensated, staged, or safely reordered. Shared agent-tool interfaces usually do not. We contribute an effect-history model that separates events in the external world from the runtime's observations of them, and a catalog of eight recurring external-effect anomalies. From the catalog we derive the boundary capabilities required to exclude each anomaly in general, and four points where black-box tool invocation alone cannot provide a general guarantee. We then ask how much of this is expressible in a widely used shared tool interface, measuring the use of the standard annotation vocabulary across 98,291 tools exposed by registered Model Context Protocol (MCP) servers. The fields are widely emitted but provide only coarse call-level hints, and none of the required capabilities is fully expressible. These results motivate reusable transactional contracts at the tool boundary.

Mon 14 SeptArtificial IntelligenceDatabasesDistributed, Parallel, and Cluster Computing
The gist
When AI agents ask external tools to do tasks during complex workflows, sometimes things go wrong behind the scenes. The authors explain how calls to these tools can appear successful but still cause problems like missing, duplicate, or leftover effects. They studied how current shared interfaces don’t fully capture important details to prevent these errors. Their work suggests better ways to track and manage these tool interactions to avoid workflow breakdowns.
Open 2609.15397v1

Avatar improves scientific workflow efficiency with AI orchestration

Avatar: Toward Autonomous End-to-End Orchestration of Scientific Workflows using LLMs

Abstract: Scientific workflow management (WMSs) systems automate execution, yet orchestrate using fixed, hand-tuned rules. LLM agents promise more autonomous orchestration, but it remains unclear where to introduce agentic reasoning, how to bound its risk, and when it actually helps. We present Avatar, an actor-based architecture comprising an orchestrator, an executor, and a provenance monitor. Each actor's decision policy is pluggable (rule-based or LLM-backed) via a single adapter-validated action catalog, so conventional and agentic control run on the same core across different WMSs. We present an implementation using the Academy framework and evaluate Avatar across three workloads. We observe that Avatar's rule mode reproduces native execution, with a single unchanged core running all three. Moreover, LLM-backed Avatar reports a reduction of compute wastage by $55\%$ and cuts GPU-busy time by $40\%$. Overall, we envision Avatar as a step toward workflow systems that reason about their own orchestration rather than follow pre-fixed rules.

Wed 9 SeptDistributed, Parallel, and Cluster ComputingMultiagent Systems
The gist
Scientific workflows automate research tasks but usually follow fixed rules. The paper presents Avatar, a system that uses AI (large language models) to decide how to run and manage these workflows more flexibly. Avatar can switch between traditional rule-based control and AI-driven decisions, reducing wasted computation and GPU busy time significantly. This work shows a way for scientific software to manage itself more intelligently instead of just following preset instructions.
Open 2609.10509v1