Papers for

customer support platform engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Progressive disclosure improves agent skills quality but adds delay

Report: Progressive Disclosure of Agent Skills

Abstract: Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost. Progressive disclosure (lazy-loading) of skills as needed may reduce operational costs, but its impact on overall latency and skill-retrieval quality remains unclear. In this report, we investigate the impact empirically and find that progressive disclosure improves skill-retrieval quality but marginally degrades overall latency.

Mon 28 SeptArtificial Intelligence
The gist
Large language model agents can do more things when given extra skills, but having too many skills makes them expensive to run. The authors studied turning on skills only when needed, called progressive disclosure, to save costs. They found this approach makes it easier for the agent to pick the right skill but slightly slows down how fast it answers. So, adding skills gradually helps quality but at a small speed cost.
Open → 2609.35692v1

Stashbird cuts memory tokens for ai agents with indexed updates

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents

Abstract: AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.

Mon 28 SeptArtificial IntelligenceMachine Learning
The gist
AI assistants need to remember past conversations accurately, including who said what and when information changes. The authors created Stashbird, a system that organizes memory with clear links to original conversations and allows easy updates or deletions. This system uses much fewer resources to process and recall information than some existing methods, while maintaining similar or better accuracy. Stashbird is tested across multiple benchmarks involving long-term conversational memory.
Open → 2609.34242v1