Papers for

customer support engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

AutoTailor improves web agent speed and accuracy by cutting unused tools

AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

Abstract: Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compact set of trajectory-derived Model Context Protocol (MCP) APIs. Offline, AutoTailor converts web trajectories into parameterized browser-automation programs, applies a Quality Filter to remove APIs with unsuitable granularity and redundant functionality, and applies a Usage Likelihood Filter to prioritize broadly useful capabilities while preserving semantic coverage. Online, Dynamic Reselection monitors task outcomes and API usage, identifies recurring coverage gaps, adds relevant candidates, and prunes persistently unused capabilities. We evaluate AutoTailor on 106 WebArena Postmill tasks. Offline filtering reduces the initial 1,283 unrefined APIs to 87, and Dynamic Reselection produces a 33-API set. With reasoning and acting (ReAct) fallback, this set achieves 90.6% correctness, compared with 87.5% for ReAct alone, while reducing average total request-token cost by 57.8% and latency by 29.4%. Without ReAct, it achieves 60.1% correctness, marginally matching the performance of unrefined set, while reducing request-token usage by 94.9%. Together, these results show that static filtering produces a compact inventory of APIs expected to support core, high-likelihood tasks, while dynamic reselection further tailors that inventory to observed user needs. This combination improves accuracy and latency while sharply reducing token usage and end-to-end cost, demonstrating the value of user-aligned capability management for efficient web agents.

Fri 11 SeptArtificial IntelligenceSoftware Engineering
The gist
Web agents use tools to interact with websites automatically, but having too many or poorly chosen tools can slow them down and waste resources. The authors created AutoTailor, which automatically picks a smaller, smarter set of tools tailored to what users actually need. It filters out unnecessary functions and adjusts the tools as it learns what works best, making the agent more accurate and faster while using fewer computing resources. This method showed better performance on tasks compared to using many unrefined tools or fallback strategies alone.
Open → 2609.13548v1

Retrieval-confidence layer detects missing context in enterprise code generation

RCL: A Retrieval-Confidence Layer for Detecting Insufficient Context in Enterprise Retrieval-Augmented Code Generation

Abstract: Retrieval-Augmented Generation (RAG) for code generation has been studied extensively on public repositories, where a model's parametric knowledge often compensates for imperfect retrieval. This breaks down in enterprise codebases, where private APIs, internal frameworks, and undocumented team conventions fall entirely outside any model's pretraining distribution. Recent work on private-library code generation shows that even oracle (perfect) retrieval does not eliminate errors, but locates failures downstream in API usage; separately, confidence-gated retrieval has been studied for open-domain question answering using model-internal confidence. Neither addresses whether retrieval itself was structurally sufficient for a private-code query before generation begins. We introduce RCL (Retrieval-Confidence Layer), a lightweight module inserted between retrieval and generation that combines a call-graph-derived structural coverage score with a novelty score estimating a query's dependence on knowledge outside the model's prior, to detect insufficient retrieval before generation occurs. When confidence falls below a calibrated threshold, RCL triggers a targeted follow-up retrieval or labels the output for human review, rather than generating silently against incomplete context. We describe RCL's architecture, formalize its scoring functions, and propose an evaluation methodology using a private-code benchmark built by injecting synthetic internal APIs into open-source Java repositories, simulating the enterprise condition without proprietary code. We report results (Section 7) comparing RCL against similarity-only retrieval on generation correctness. Our position is that retrieval sufficiency, assessed structurally rather than from model-internal confidence, is a distinct and currently underaddressed signal for building safer code-generation systems in private, enterprise settings.

Thu 10 SeptSoftware Engineering
The gist
Generating code with AI often relies on retrieving related code snippets, but in private company codebases, this retrieval misses important internal or undocumented parts. The authors propose a new method called RCL that checks whether the retrieved code covers enough relevant parts before the AI tries to write new code. If it detects missing pieces, it can call for more searching or ask for human help instead of risking wrong code. They tested RCL using a simulation with injected internal APIs to mimic private code environments.
Open → 2609.11023v1