Papers for

automated customer service teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Large reasoning models get better uncertainty estimates without internals

Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models

Abstract: While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.

Mon 28 SeptArtificial IntelligenceComputation and LanguageMachine Learning
The gist
Large reasoning models (LRMs) can be too confident in their answers, which is risky if we can't see their internal scores. The authors find that usual ways to guess how uncertain these models are don't work well because the models give less varied answers after fine-tuning. They introduce a new way to 'jailbreak' the model's answers to get more realistic uncertainty estimates without needing internal access. Their method, called J4U, shows better uncertainty predictions across several datasets and models.
Open → 2609.35350v1

Fisher-informed method improves stability of feedback learning in large language models

Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs

Abstract: Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.

Sun 27 SeptMachine Learning
The gist
Training large language models to learn from their own outputs can be unstable and cause performance to collapse. The authors propose a method called FIRE that helps the model better handle feedback by separating how feedback moves the model from how far it moves. This method recalibrates training signals based on whether answers are correct or incorrect, using information from the model’s internal statistics. Their experiments show that FIRE makes training more stable while keeping the model’s performance strong.
Open → 2609.34009v1

Language model agents design their own evaluators to improve task success

Self-Designed Evaluators and Warm Memory for Long-Horizon Agents

Abstract: A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite's evaluator gives a small, consistent gain, and using them to warm up ACE's memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.

Sun 27 SeptArtificial Intelligence
The gist
When a language model agent works on many tasks in a row without clear success signals, it can't tell if it did well or try again safely. The researchers created SelfSuite, where the agent uses its own model to build a set of small tests and judges to check its work. This system helps the agent decide when to retry tasks and learn from past results without needing labels from experts. SelfSuite performs better than agents without such feedback and is competitive with methods that rely on expert labels.
Open → 2609.33717v1

Agentic systems improve reasoning fairness and efficiency

Agentic Multi-Turn Reasoning: A Fairness Approach

Abstract: Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $Φ$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.

Sun 27 SeptArtificial IntelligenceMachine Learning
The gist
Solving complex problems often requires reasoning through many steps, but teaching AI to do this well is tricky when feedback only comes after all steps are done. It’s even harder when the data mainly shows common reasoning paths, making rare but important approaches overlooked. The authors developed a new method called Fair Multi-Level Preference Optimization that helps AI models learn better by handling this long-step challenge and data imbalance fairly. Their approach showed improved performance on benchmarks designed to test step-by-step AI reasoning.
Open → 2609.33323v1

Verification layer improves large language model agent termination

Verification as an Architectural Layer for LLM Agents: A V-Model Design, and a Pilot Study of Its Deterministic Core

Abstract: Large language model (LLM) agents built on the ReAct pattern concentrate four responsibilities in one model: selecting a strategy, choosing each action, formatting it, and judging whether the result is adequate. Nothing outside the generative loop can reject its output, so an agent that cannot make progress does not report failure; it runs until an external budget stops it. We propose treating verification as an architectural layer by adapting the V-model from software engineering: specification levels descend from requirements to individual steps, each level is paired with a dedicated verifier, a deterministic controller enforces every verdict, and only verification outcomes write to memory, so a rejection localizes the level that introduced the fault and an agent halts by declining rather than by exhaustion. Each verifier separates a zero-cost deterministic \emph{gate} from an optional LLM \emph{judge}, so the contribution and cost of each can be measured independently. We report a pilot implementing the acceptance- and unit-level verifier pairs, comparing five configurations that share one executor, tool set, and scorer and differ only in verification, on the four-hop stratum of MuSiQue with an 8B-parameter backbone. Across 47 executions, the two unverified configurations answered none of ten questions, every run ending at a step cap or provider token limit; the verified configuration without a planner answered eight and abstained on the rest. Deterministic gates produced eight of the nine observed corrections at zero marginal cost, and planning degraded performance once verification was present. These results characterize termination behavior, not accuracy at scale; we outline a twelve-month plan to complete and evaluate the full architecture, including the integration-level pair the pilot omits.

Fri 25 SeptSoftware EngineeringArtificial Intelligence
The gist
Large language model agents that try to solve complex tasks often keep going endlessly if they can’t finish properly. The authors suggest adding a verification step that checks each part of the process and stops the agent early if there’s a problem. This makes the agent more reliable and avoids wasted effort. They tested a small version showing it helped the agent give fewer but more certain answers and stop correctly instead of running out of resources.
Open → 2609.31937v1

Ai agents struggle to reject misleading user suggestions

XYEval: Agents say yes to bad advice

Abstract: Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where a person asks about their attempted solution rather than their actual problem. We extend prior sycophancy evaluation to the XY problem in agentic settings, evaluating whether agents can resist plausible but misleading suggestions from users and communicate their reasoning. We introduce XYEval, a meta-evaluation framework that can transform an existing benchmark into an XY problem evaluation. We evaluate five models across six diverse benchmark suites. Agents suffer large XY drops under XY mutation across benchmarks, with relative drops reaching up to 46.7%. With $τ^2$-bench, we further show that agent performance drops more when encountering a pedantic user who requires detailed explanations before approving a better solution. Our findings suggest that current agents lack the ability to effectively reason and communicate when facing misleading suggestions. A simple system instruction baseline that encourages awareness of XY problems only offers partial mitigation. Extensive trace analyses provide behavioral insights into how and why these XY drops occur across execution trajectories. Our results show that mitigating the XY problem remains challenging, requiring agents to both recognize user misdirection and clearly communicate the underlying problem.

Sun 20 SeptComputation and Language
The gist
AI agents often fail to recognize when users give bad or confusing advice, leading to poor decisions. The authors created a way to test how well different AI models handle these tricky situations, called the XY problem. They found many agents perform much worse when users give misleading information and that simple instructions to be cautious only partly help. This shows current AI agents need better reasoning and communication skills to handle misleading user input.
Open → 2609.23939v1