Papers for

conversational ai teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

New method improves language model learning from verbal feedback

Inductive Feedback for Mixed-Policy Distillation

Abstract: Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences that the feedback did not motivate, while leaving much of the feedback's guidance unused. We find that both problems come from the standard on-policy distillation objective, specifically the divergence it minimizes and the distribution it uses as its target. Our proposed method addresses both limitations. First, to isolate the information conveyed by the feedback from the teacher's inherent preferences, we treat verbal feedback as evidence for or against the hypothesis that a particular token comes next at a given prefix. We then adopt a probabilistic confirmation framework which uniquely determines an ordering over the vocabulary based on the teacher's predictions before and after it receives feedback. Using a confirmation score consistent with this ordering, we construct a target distribution within a trust region of the student. Second, to learn from guidance that student rollouts can leave unused, we derive a simple shared-rollout estimator of a symmetric divergence between the student and target distributions over rollouts, reusing student and feedback-conditioned teacher rollouts in both directions through importance weighting. Empirical evaluations show that our method outperforms the common on-policy distillation recipe and a recent contrastive variant on knowledge-based and agentic benchmarks.

Mon 28 SeptMachine Learning
The gist
Sometimes language models can be taught better by using verbal feedback that points out mistakes and suggests fixes. The authors found that current ways of using this feedback can accidentally transfer the teacher model's own biases or ignore much of the feedback. They developed a new approach that treats feedback as clues for what the next word should be and carefully updates the student model’s learning. Their method also makes better use of the feedback even when the student model doesn’t immediately apply it. Tests show this approach works better on tasks involving knowledge and agent actions.
Open → 2609.35390v1

LLM agents struggle to plan and update personalized tool tasks over time

PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents

Abstract: Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive execution, leaving unclear whether agents can formulate, execute, and revise an explicit plan while preserving user preferences throughout long-term interaction. To address this gap, we introduce \textbf{PDEU-Bench} (\textbf{P}ersonalized plan \textbf{D}efinition, plan \textbf{E}xecution, and plan \textbf{U}pdate \textbf{Bench}mark), a benchmark for evaluating the complete planning lifecycle of personalized tool-using agents. PDEU-Bench comprises 214 long-horizon interaction tasks spanning 12 everyday domains and 94 tools, with stage-specific assessments of preference adherence and plan quality. Extensive evaluations of 15 representative open-source and closed-source LLMs reveal a pronounced gap between local tool execution and dynamic planning: LLMs can often instantiate preferences in individual calls, yet struggle to construct coherent plan definition and plan update. We further evaluate mainstream personalization and memory-augmentation methods. Although these methods improve particular stages, none of the evaluated methods reliably propagates user preferences throughout the complete lifecycle, and their gains frequently fail to transfer to subsequent execution. Fine-grained error analysis further reveals that preference omissions and conflicts persist throughout the planning lifecycle, highlighting the need for future research to parameterize LLMs with preference-aware information retrieval and memory capabilities. We provide the relevant code and data in the appendix to support future research.

Mon 28 SeptArtificial Intelligence
The gist
People want AI agents that can carry out multi-step tasks while remembering what users prefer over time. The authors created a new test called PDEU-Bench to check how well language model agents can make, follow, and adjust plans while respecting personal preferences. They found that while models handle individual tool calls okay, they have trouble with overall planning and updating plans. Current personalization methods help somewhat but don’t fully solve these problems yet.
Open → 2609.34930v1

Method improves diversity of AI text generation without trade offs

Improving the Diversity of LLM Outputs without a Trade-off

Abstract: We propose DAST (Diversifying Arithmetic Sampling with TokenTour), a method that increases the diversity of LLM outputs without any change to the marginal distribution and with negligible generation-time overhead (a few microseconds). We observe that token IDs are often arranged in a meaningless order and reassign them so that tokens with similar meanings appear consecutively. This can be done in advance in a few hundred seconds per model, and the resulting order can be reused for all subsequent generations. By combining this order with arithmetic sampling (or quasi-Monte Carlo methods), we make similar tokens less likely to be generated across runs while preserving the distribution. Our method not only produces qualitatively good ideas but also significantly improves performance on the downstream task of ProtoQA.

Sun 27 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Language models sometimes produce similar or repetitive responses, limiting the variety of ideas they generate. The authors propose a method called DAST that rearranges how tokens (words or pieces of words) are ordered internally so that similar meanings appear together. This change helps make similar tokens less likely to be picked repeatedly, increasing output diversity without slowing down generation or changing what the model thinks is likely overall. Their approach improves results on a question-answering task, showing it can produce a wider range of good answers.
Open → 2609.33038v1

Language models improve counting by thinking and revisiting evidence

Counting on Thinking: Tracing Evidence Integration in Language Models

Abstract: Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operation humans and animals perform automatically. We ask why this requires thinking in LLMs. Evidence integration has long been used in psychology and neuroscience to probe decision-making. Our evidence-integration task presents one letter per conversational turn and asks which of two target letters appeared more often. A running count difference solves the task optimally by weighting every letter equally; tokens at each turn could represent and update this difference. Direct responses instead weighted evidence unevenly, with strong recency effects, and assigned less probability to the correct answer as difficulty increased. Thinking improved performance and made integration weights nearly uniform, yet final-query attention remained concentrated on the sequence ends in both modes. Reasoning trajectories showed models revisiting input, recounting letters, and checking intermediate counts that informed the answer, suggesting that thinking constructs the accumulated count that direct responses lack rather than reading out one already formed. Reasoning-token costs grew with the number of letters far more than with coherence. Outcome feedback did not bring this computation into direct responses: under in-context reinforcement learning (ICRL), performance deteriorated over repeated games and recency effects strengthened, yet models grew more confident. Humans and animals amortize such computations into automatic processes, whereas current LLMs still pay for them with thinking on every trial. Which operations learning can make directly available remains central to how future models allocate computation.

Sat 26 SeptMachine Learning
The gist
Counting which letter appears more in a sequence is easy for humans but hard for large language models (LLMs). The authors studied why LLMs struggle by giving them tasks where they must keep track of letters one at a time. They found that without extra 'thinking' steps, LLMs guess based more on recent letters and perform worse as tasks get harder. When asked to think longer, the models do better by reviewing and updating counts more evenly. However, unlike humans, LLMs still have to spend effort on counting each time instead of doing it automatically.
Open → 2609.32932v1

Moscopt improves llm agents by mixing and choosing skills dynamically

MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents

Abstract: Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template---missing the synergy among multiple complementary strategies. We propose MOSCOPT, a text-native, parameter-free algorithm that jointly optimizes a pool of $N$ skills and a gating skill $G$ that dynamically selects $K$ skills per step. To effectively optimize the skills, we build the EditAdam with internally maintained dual states. Through the three-phase interleaved updates with EditAdam, the system monotonically improves without gradient or parameter tuning. Extensive experiments and detailed ablations across 5 benchmarks and 3 target LLMs demonstrate that MOSCOPT consistently outperforms all baselines, and confirm that both the mixture-of-skills architecture with selective activation and the collective evolution with three-phase interleaving are essential to its superior performance. Code is released https://github.com/zhangzhenyu13/SummerClaw/tree/master/summerclaw/agent_trainer/algorithms/moscopt.

Sun 13 SeptArtificial IntelligenceComputation and Language
The gist
LLM agents use language prompts called skills to perform tasks, but existing approaches optimize only one skill at a time, missing benefits from combining multiple strategies. The authors propose MOSCOPT, a method that optimizes a group of skills together and learns when to activate which skills dynamically. They introduce a new optimization technique called EditAdam to improve these skills without needing parameter tuning. Experiments show MOSCOPT performs better than previous methods by using a mix of skills and updating them collectively in an organized way.
Open → 2609.14399v1