Papers for

customer service platform developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Reliable multi-turn business agent evaluation with modular LLM judges

Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents

Abstract: Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to either agent capability or missing business knowledge before they are actionable. We present an integrated methodology spanning evaluation specification, modular LLM judges, intent-preserving user simulation, and human-in-the-loop governance. The specification defines conversation-level end states and actionable failure ownership. Atomic judges share versioned evidence and feed an explicit aggregation graph. The simulator is released only after task-preservation and stability checks. Independent human audits estimate measurement fidelity, renew tiered reference sets, and route disagreements to label correction, guideline revision, or judge improvement. Production studies show that system-level fidelity improved across repeated audits, that human reviewers and automated judges improved together under the shared feedback loop, and that their combined workflow had the strongest descriptive performance in both reported task-completion settings. Because the studies are observational and the human reference itself required revision, these findings demonstrate operational usefulness rather than causal or universal superiority. The contribution is a practical framework for making multi-turn agent measurement reliable, actionable, and maintainable as the evaluated system and its evidence evolve.

Sun 27 SeptArtificial Intelligence
The gist
Evaluating AI business agents over multiple conversation turns is tricky because success depends on complex facts and steps that unfold over time. The authors developed a method that uses small, specialized AI judges combined with simulations and human checks to measure how well these agents perform. Their system helps figure out if problems come from the AI’s abilities or missing business knowledge, making fixes easier. Tests in real scenarios showed this method improves judging accuracy and helps humans and AI judges learn from each other.
Open → 2609.33955v1

Generative AI agents need resilience and care over time

Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

Abstract: Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.

Wed 9 SeptArtificial IntelligenceHuman-Computer InteractionMultiagent Systems
The gist
Getting AI to finish a task isn’t enough for long-term use, especially in places like healthcare. The authors found that AI agents need to bounce back from problems and work well with people who depend on them. They studied how AI handles challenges over many interactions and saw that agents change how they recover and how they pay attention to others. The paper highlights important trade-offs when deploying AI to help people in ongoing, real-world workflows.
Open → 2609.10724v1