Papers for

software development tool builders

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Reward aligned weighting improves student model training accuracy

Reward-Aligned Reweighting for On-Policy Distillation

Abstract: On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R$^{2}$-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R$^{2}$-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.

Mon 28 SeptMachine Learning
The gist
Teaching a smaller language model to think like a bigger, smarter one usually treats all parts of an answer as equally important. The authors found this is not always the best way because some steps are more important to get right for finishing a good answer. They created a way to focus the teaching on the parts that really matter, based on how well the bigger model and smaller model agree on the final result. This method helped smaller models do better on math problems and coding tasks.
Open → 2609.35517v1

RepoAtlas helps coding agents find and fix code problems faster

RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

Abstract: Large language model (LLM)-powered coding agents have made rapid progress in automating software engineering tasks, yet repository-level issue resolution remains challenging. Beyond generating a plausible patch, an agent must localize relevant code across interdependent files and maintain repository context that is both sufficient and focused. Code graphs expose non-local relations, but linear text interfaces obscure their topology; rendering the full repository graph yields visual representations that are too dense to perceive reliably, whereas a one-shot local view becomes stale as exploration proceeds. We present \textbf{RepoAtlas}, a training-free module that maintains evolving multimodal repository views through a \emph{select--project--refresh} loop over a repository code graph. RepoAtlas combines evidence from the issue with the agent's current exploration state to select a task-relevant region under a fixed budget, projects the selected structure into complementary visual and textual representations, and refreshes the view when changes in the exploration state render it outdated. We evaluate RepoAtlas on SWE-bench Verified, where it improves the resolve rate by 2.4 points while reducing input tokens and model calls by 5.8\% and 7.8\% on average, relative to the strongest multimodal graph baseline, with consistent gains across three models of different families and scales.

Tue 15 SeptSoftware EngineeringArtificial Intelligence
The gist
Fixing bugs in large software projects is hard because related code is spread across many files. The authors present RepoAtlas, a tool that helps AI coding assistants keep track of the most relevant parts of a project by combining visual and text views of code dependencies. RepoAtlas updates these views as the AI explores the code, so it stays focused on the important pieces. This approach improved bug-fixing success rates and made the AI more efficient in tests.
Open → 2609.16936v1