Papers for

chatbot development teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Continuous context management reduces language model context size with tradeoffs

Continuous Context Management

Abstract: Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.

Mon 28 SeptArtificial Intelligence
The gist
Language model agents usually keep their full conversation history, which makes the context large and slow over time. The authors study a method called Continuous Context Management (CCM) that shrinks this history after every step, keeping only a summary and the latest observation. This approach cuts down how much information the model processes but can also reduce task success depending on the model used. They improve CCM by training models with extra feedback so CCM performs better, making it a promising way to keep language models efficient with less memory.
Open → 2609.35540v1

Simulation improves customer service AI agents without risking customers

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

Abstract: Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.

Thu 24 SeptArtificial IntelligenceComputation and Language
The gist
Customer service chatbots need to handle complex requests carefully, especially in sensitive industries. The authors created a way to test these chatbots using fake customers and simulated tools, so they don't cause problems for real users. By doing this, they could try many versions and improve the chatbot's quality before letting it interact with actual customers. Their method helped increase positive customer feedback and the use of self-service options without harming customer trust.
Open → 2609.30137v1