Papers for

ai application developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Capfield-opd enables smooth control of multiple capabilities in ai models

CapField-OPD: Learning Continuous Capability Fields via Joint-Anchored Multi-Teacher On-Policy Distillation for Flow Models

Abstract: Reward-specialized post-training produces strong experts for flow-based generative models, while multi-teacher on-policy distillation (OPD) consolidates their capabilities into a single student. Existing methods, however, route each prompt to a single teacher according to its semantic category, implicitly binding the desired capability to prompt content. This coupling makes capability invocation vulnerable to prompt perturbations and prevents users from explicitly adjusting the strength of the desired capability at inference time. In this work, we introduce CapField-OPD, an OPD framework that integrates multiple teachers into a continuous capability field through explicit capability coordinates. We use teacher models as anchors to construct this field, with the coordinates determining how their outputs are combined. Each capability configuration thus receives a unique supervision target, and capability control no longer depends on prompt semantics. Since the training anchors may not be optimal at inference time, we further profile the learned field on a small calibration set. The coordinate with the highest mean reward serves as the recommended default, while coordinates that are frequently optimal offer a promising candidate set for test-time scaling. Extensive experiments on compositional generation, text rendering, and visual aesthetics demonstrate that CapField-OPD consolidates multiple specialized teachers into a single student while preserving or surpassing their performance, reliably invokes the desired capabilities under semantics-preserving prompt variations, and supports continuous capability control and coordinate-based test-time scaling.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
AI models often learn specialized skills from different teachers, but switching between these skills usually depends on specific words or prompts, which can cause errors if the prompts change slightly. This paper introduces a new method that combines multiple specialized AI teachers into one model, allowing users to smoothly control and adjust the model's skills by using coordinates instead of relying on exact prompt wording. This approach makes the model more reliable and flexible when performing tasks like creating images or rendering text.
Open → 2609.34658v1

Initial user input guides autonomous deep research outcomes clearly

What Happens During Autonomous Deep Research After the User Steps Away?

Abstract: In autonomous deep research, a user provides a task and relevant background, then leaves the agent to conduct an extended investigation without further human intervention. We study how this initial user information is reflected in intermediate actions and how these actions relate to final recommendations. We introduce DRaligned, a counterfactual behavioral evaluation framework built on PDR-Bench. By varying one task-relevant user factor while keeping the remaining context fixed, we compare acquisition requests, working drafts, and final reports. Source-grounded extraction, blinded local judgments, and deterministic aggregation yield coarse directional measurements while leaving ambiguous cases unresolved. Our experiments show that strong user-specific delivery can emerge from a largely shared research process: agents investigate similar broad questions but allocate requests differently, and final recommendations distinguish user conditions more clearly than explicit requests do. Reports can also integrate user factors that were not jointly visible during acquisition. In readable draft-to-report comparisons, recommendations often retain their coarse user-specific direction despite substantial rewriting. Final directional differences recur across tested agent models, execution harnesses, and evaluator models, even as execution paths vary. These findings describe how initial user information shapes autonomous research and clarify the relationship between the process an agent follows and the recommendations it delivers.

Sun 27 SeptArtificial Intelligence
The gist
When users give an AI system a research task and then step away, the system keeps working on its own to complete the job. This paper studies how the initial instructions influence what the AI investigates and the final advice it provides. The authors found that even though the AI explores similar topics, it tailors its requests and final reports based on who gave the instructions. The research shows that the process and final answers stay aligned with the initial user’s needs, even if some parts change during the investigation.
Open → 2609.33509v1

Calibrating reasoning models improves confidence estimates greatly

Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

Abstract: Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.

Sun 27 SeptComputation and Language
The gist
Models that answer questions often say how confident they are, but this confidence is usually too high and doesn't match the true chance of being right. The authors found a better way to measure confidence by checking the model’s internal reasoning steps, which gives more accurate confidence levels. They also developed a method to teach the model to trust this better confidence measure, improving its predictions without needing complex new training techniques. This helps models not just pick answers, but also know how sure they should be about them.
Open → 2609.33290v1

Large language models improve reasoning by optimizing skills prompts and routing

Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs

Abstract: Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design space: they either optimize a monolithic prompt, or separately induce and refine skills from model traces. As a result, they lack a principled mechanism for deciding which component should be updated when failures occur, and they rarely optimize prompts, skills, and skill-use policies in a unified framework. We propose SPARO (Skill, Prompt, And Routing Optimization), a framework that jointly optimizes task instructions, reusable skill blocks, and routing rules. It performs controlled counterfactual evaluations, converts examples' effects into a probabilistic responsibility distribution over prompt, skill, and routing components, samples one component from that distribution, and applies the corresponding targeted mutation. This design moves language-program optimization beyond global prompt rewriting toward structured, reusable, and selectively activated task knowledge. Across five benchmarks and five worker models, SPARO consistently outperforms both prompt-centered and skill-centered optimization baselines. These results suggest that effective language-program optimization depends not only on discovering useful task knowledge, but also on deciding where that knowledge should be stored and when it should be activated.

Sat 26 SeptArtificial Intelligence
The gist
Language models can solve many tasks, but how you tell them what to do and which parts of their knowledge to use is very important. The researchers introduce a method called SPARO that improves language model programs by adjusting task instructions, reusable skills, and how the model chooses which skills to use. This method figures out which part to change when the model makes mistakes and updates that part to do better next time. Their tests show this approach works better than just changing the instructions or skills alone.
Open → 2609.32492v1

Large language models improve skills by fixing mistakes locally

A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

Abstract: Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify where productive problem solving begins to break down, rather than reflect coarsely over the entire failure. Based on this insight, we propose SkillPivot, a deviation-point-guided framework for skill self-evolution. SkillPivot detects the transition from a useful prefix to an erroneous suffix using execution validity, goal progress, and action diversity. A stronger teacher then continues from the same prefix and produces a successful alternative under the same interaction history. By contrasting the student's failed suffix with the teacher's successful suffix, SkillPivot generates localized skill updates while preserving already effective guidance. Experiments on ToolQA, LogicBench, and WildClawBench show that SkillPivot consistently outperforms competing skill-evolution methods, improves multiple agent models, and produces compact, transferable skill updates.

Thu 24 SeptArtificial Intelligence
The gist
Complex tasks can often be done in many right ways, so forcing a language model to copy one fixed correct solution doesn’t help it learn well. The authors found that even when a model’s attempt fails, it often makes good progress before going off track. They developed SkillPivot, which identifies where the model’s reasoning starts to fail and then uses a stronger model to fix just that part. This focused approach helps the model improve without losing what it already does correctly.
Open → 2609.29154v1

Policy distillation improves reinforcement learning outcomes beyond initial accuracy

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

Abstract: Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students initialized with OPD reach higher final performance than those trained with direct RL or supervised fine-tuning followed by RL. This advantage can emerge even when OPD produces little immediate improvement in accuracy. Pre-RL Pass@k does not fully explain the benefit: similar or even higher values do not necessarily lead to better performance after RL. Behavioral analyses point to alignment with the teacher's distribution beyond top-1 agreement as a possible explanation. Such alignment may favor higher-quality reasoning paths while retaining alternatives that RL can further refine using outcome feedback. We further examine how trajectory sources and divergence objectives affect the value of distillation for subsequent RL. Standard reverse-KL OPD performs better before RL, but forward-KL OPD overtakes it afterward; with teacher-generated distillation trajectories, reverse KL remains ahead at both stages. These findings suggest that the preferred distillation objective depends on both the trajectory source and the training that follows. Our results support evaluating OPD as preparation for RL and selecting distillation choices by the performance achieved after subsequent training.

Wed 23 SeptMachine Learning
The gist
Reinforcement learning (RL) helps machines learn by trial and error, but it matters which starting point you choose. The authors studied a way to prepare for RL called on-policy distillation (OPD), where a student learns from a teacher’s behaviors before actual RL training. They found that starting with OPD can lead to better final performance, even if it doesn’t improve the student’s accuracy right away. This improvement might come from the student better matching the teacher’s range of good behaviors, giving RL more options to improve upon.
Open → 2609.28145v1

Planned test-time scaling improves reasoning task performance with coordinated problem solving

Planned Test-Time Scaling with Coordinated Reasoning Paths

Abstract: Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address this limitation, we propose Planned Test-Time Scaling (PTTS), which replaces independent sampling with a coordinated joint policy: a planner generates a solution outline for each branch, steering the branches toward distinct reasoning paths, and an executor produces a full solution conditioned on each outline. Formally, we show that PTTS strictly generalizes repeated sampling and, in a stylized setting, provably promotes coverage of complementary reasoning modes and yields better pass@k scaling. We instantiate PTTS on top of strong reasoning models, keeping them fixed as executors while replacing repeated sampling with PTTS inference to further enhance test-time scaling. Concretely, we develop two variants: PTTS-ZS prompts a model to jointly generate outlines for all branches in a single autoregressive pass, while PTTS-RL directly optimizes the planner against the pass@k reward using truncated execution rollouts for efficient training and a sharper reward signal. Across five mathematical reasoning benchmarks with Qwen3-1.7B and 4B, PTTS-ZS improves pass@64 over repeated sampling by up to 6.7 points, while PTTS-RL further increases the gain to up to 13.4 points. Further analysis indicates that broader coverage of distinct reasoning paths contributes to these gains. Overall, PTTS provides a general framework for improving test-time scaling by coordinating reasoning branches, with zero-shot and trainable instantiations that yield substantial performance gains.

Wed 23 SeptComputation and LanguageArtificial Intelligence
The gist
When computers try to solve hard problems, they often try many guesses at once to find answers. The authors show that making these guesses work together, instead of independently, helps the computer explore different kinds of solutions better. They created a new method called Planned Test-Time Scaling (PTTS) that first plans different outlines for each guess and then completes each one. This approach improves success rates on math reasoning tests by encouraging diverse ways of thinking. It works both without extra training and with training to get even better results.
Open → 2609.27374v1

Vision language models lose accuracy when image layout changes

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as visual question answering, yet small image resizes can turn correct answers into errors. We investigate whether changes in visual configuration, such as image tiling and token arrangement, contribute to this instability. Across seven checkpoints and four benchmarks, equally small resizes cause more correctness flips when they switch configurations. Surprisingly, in over half of these cases, models answer the question incorrectly but can still read the correct answer when told what to read. Furthermore, attention interventions in LLaVA-NeXT suggest that configuration changes can weaken the use of readable information during answering. We therefore guide models using field cues and their own transcriptions. With annotation assistance, these forms of guidance together correct 97.2% of errors with readable information. These findings show that configuration changes can affect how models use information they can still read.

Tue 22 SeptComputer Vision and Pattern Recognition
The gist
Vision-language models can answer questions about images well, but small changes in how the image is arranged can cause these models to give wrong answers. The authors found that even when the model reads the correct answer in the image, it sometimes fails to use that information correctly because the way the image is shown has changed. They showed that by helping the model focus on the right parts of the image or text, most of these errors can be fixed. This means changing the image layout affects how well the model uses the information it can actually see.
Open → 2609.25770v1

SkillAA improves AI skill updating with precise graph-based editing

SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback

Abstract: External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation. We introduce SkillAA (Skill Abductive Attribution), a structured skill-optimization framework for frozen language models. It represents skill applicability, execution, and composition in a unified graph, allowing the same structure to support skill selection, attribution-guided repair, and update validation. SkillAA contrasts successful and failed executions to route candidate repairs to specific graph objects, updates only the selected local structure, and uses Local and Big Gates to screen candidate changes before commitment. With gpt-5.6-sol, SkillAA reaches 81.5%, 66.7%, and 91.2% on SearchQA, LiveMath, and DocVQA, respectively, and attains the highest observed mean in every main setting. These results support the utility of attribution-guided graph editing and graph-scoped validation.

Thu 17 SeptArtificial Intelligence
The gist
AI systems often use fixed skills to perform tasks, but fixing errors in those skills can be tricky and inefficient. The researchers created SkillAA, a method that uses a detailed graph to represent skills and how they connect, allowing the system to locate the exact part that caused a failure and update only that part. This makes AI skill updates more accurate and reliable. Testing with a powerful language model showed SkillAA achieves high success rates on several question-answering tasks.
Open → 2609.20455v1

Visual grounding improves with confidence awareness to reduce hallucinations

SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation

Abstract: Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with preferences such as DPO variants, which teach which answer is preferred but not when the model's own answer is unreliable. We argue that calibrated self assessment is the missing signal. We introduce Savor, a training framework that (i) augments the output schema with token and answer confidence, (ii) optimises the policy with a Group Relative Policy Optimisation (GRPO) objective that penalises calibration error and poor abstention decisions, and (iii) uses the learned confidence at inference time to revisit visual evidence only when the model is uncertain. Experiments on POPE, HallusionBench, AMBER and MMHal-Bench across two recent backbones (InternVL3-8B and Qwen3-VL-8B) show that Savor reduces hallucination while preserving general capability on MME and MMBench, with lower Expected Calibration Error than DPO and decoding baselines.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Multimodal large language models sometimes make confident but incorrect claims about images, describing things that aren’t really there. The authors found that teaching these models to be aware of their own confidence helps them avoid making false statements. They created a training method called Savor that lets the model judge when it is uncertain and look again at the image before answering. This approach reduces mistakes while keeping the model’s overall ability to understand images and text.
Open → 2609.16601v1