Papers for

ai system trainers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Rubric response theory improves reward scoring from partial rubric feedback

Rubric Rewards from Item Response Theory

Abstract: Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.

Mon 28 SeptComputation and Language
The gist
Many tasks for AI, like grading language answers, don’t have a single correct answer. Instead, experts use rubrics with criteria to judge quality, but combining those into a single score can be tricky and inefficient. The authors propose Rubric Response Theory (RRT), a new way to turn rubric judgments into a more accurate and adaptive reward score by modeling each criterion’s difficulty and importance. This method reduces how many judgments are needed while improving scoring compared to older methods.
Open → 2609.35646v1

Motion language evaluators struggle to recognize structure and mirror actions

Can Motion-Language Models Ground Structure? STRIDE for Evaluating the Evaluators

Abstract: Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evaluate the ability of evaluators to track temporal order, mirror reflection, and action identity. STRIDE comprises $5{,}869$ triples, each consisting of a motion, its original caption, and a perturbed caption, spanning both short and long descriptions. We likelihood-balance caption pairs to reduce text-only bias and estimate each evaluator's caption preference under unrelated motions to measure the discrimination gain from matched motions relative to this baseline. Our experiments reveal weak structural grounding and severe deficits in mirror sensitivity among the audited evaluators, which commonly used evaluation protocols fail to expose. To understand why these limitations go undetected in standard tests, we examine the evaluators more closely. We find that text-only priors alone can solve naive perturbation tests on existing datasets, while common retrieval and distributional metrics barely respond to structural corruption introduced by mirroring ground-truth motions. These findings suggest a natural intervention: structural hard negatives. Our experiments show that a simple modification to contrastive learning substantially improves performance on temporal order and mirror reflection. The benchmark and code will be released.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Motion-language models are tools that match actions with descriptions, but the systems used to judge how well they do this often miss important details. The authors created a new test called STRIDE that checks if these judges can tell the right order of movements, recognize when an action is mirrored, and identify the action itself. They found that these evaluators often fail to notice mirrored actions and don't properly understand the structure of the descriptions. They also discovered that many current tests can be passed without truly understanding the action, and showed a way to improve models by training them with harder examples.
Open → 2609.32462v1

Persistent negatives stabilize reward training in black-box policy learning

Persistent Negatives for Adversarial Black-Box On-Policy Distillation

Abstract: Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent-negative adversarial distillation, a live-pool method that replaces a fraction of each discriminator batch with historical, prompt-matched teacher--student comparisons. Under matched discriminator compute, historical comparisons train the discriminator, while GRPO remains on-policy with fresh student responses. Our analysis identifies the Bayes-optimal reward as a teacher-to-negative log-density ratio and, under explicit assumptions, shows how persistent negatives anchor the discriminator and reduce reward-estimation MSE relative to fresh-negative training. Across two student families, three judges, and four judged-chat benchmarks, persistent-negative adversarial distillation consistently improves performance over current methods at matched discriminator compute. It also yields smoother fresh-policy discriminator trajectories, with fewer below-chance dips. These findings identify the discriminator's negative distribution as an important design axis in black-box on-policy distillation.

Fri 25 SeptComputation and LanguageArtificial Intelligence
The gist
Training AI systems by learning from their own generated responses can be tricky when only sample answers are available, not detailed probability info. The paper studies a way to improve this by using a technique that compares current AI responses to both recent and past examples, preventing the training from chasing a moving target. This helps the system learn more accurately and consistently. The authors show that using a mix of new and historical comparisons improves performance and stability across various tests.
Open → 2609.30864v1

Weight decay controls delayed learning and generalization in neural nets

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

Abstract: In grokking an early fit to the training data separates from a much later improvement in generalization. During this delay, training can move from a fixed neural tangent kernel (NTK) regime to one in which task-relevant kernel eigendirections continue to evolve. We provide a quantitative theory for how this transition from lazy to rich learning can produce delayed generalization. For homogeneous networks trained with squared loss and $L_2$ weight decay, we show that a finite residual remains after memorization, with larger residual fractions in target components associated with smaller NTK eigenvalues. These residuals feed back into the dynamics of the NTK itself, and projecting the resulting dynamics onto task-relevant spectral directions yields a reduced system in which residual-driven kernel growth competes with weight decay. This system predicts that the grokking timescale is controlled by the product of learning rate and weight decay, that feature learning slows logarithmically near a critical decay above which task-aligned NTK structure can no longer support generalization, and that stronger decay can prevent fitting altogether. We test these predictions in modular addition. In a homogeneous MLP, task-aligned Fourier structure continues to emerge in the NTK after training accuracy has saturated, and an 84$\times$90-grid of trained networks across varying learning rate and weight decay recovers the predicted phase geometry and inverse-product scaling of the generalization time with learning rate and weight decay. A one-block Transformer shows similar macroscopic phase structure in a 42$\times$45-grid, as well as the same transition-time scaling despite violating exact homogeneity. Together, these results provide a mechanistic derivation connecting post-fit feature learning to both the onset of generalization and its phase structure in the learning rate and weight decay plane.

Tue 22 SeptMachine LearningArtificial Intelligence
The gist
Sometimes neural networks learn to fit training data quickly but only get better at generalizing to new data much later, a phenomenon called grokking. The authors explain this delay by showing how weight decay, a common training technique, causes the network to slowly update feature representations after initially memorizing the training set. They provide a mathematical model predicting how learning rate and weight decay combine to influence the timing and success of this transition from simple memorization to richer understanding. The theory is supported by experiments on modular addition tasks using multilayer perceptrons and transformers. This helps understand why and when neural networks start to truly generalize during training.
Open → 2609.26679v1

Legal reward models improve grounded reasoning and abstention in law ai

Building Legal Reward Models for Grounding and Abstention

Abstract: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (LRB), a benchmark for evaluating grounded legal generation under noisy and insufficient retrieval conditions. Across general and legal contextual evaluation, we find that contextual DPO improves grounded evaluation, but performance is sensitive to preference-data construction. Length-balanced augmentation substantially improves grounded legal evaluation, with the strongest configuration combining length-balanced legal and general contextual preference data and improving performance by up to $\mathbf{+25.6}$pp over baseline. We further find evidence of cross-jurisdiction transfer: models contextually refined primarily on Victorian criminal-law data improve grounded evaluation on external US legal benchmarks, including a $\mathbf{+16.2}$pp improvement on \textsc{Housing Statute QA}. Together, these results provide a reproducible foundation for constructing and evaluating grounded legal reward models in retrieval-augmented settings.

Sun 13 SeptComputation and LanguageArtificial Intelligence
The gist
Computers used in legal work need to base their answers on real evidence and say "I don't know" when they lack good information. The authors created new ways to train and test models that judge if these AI answers are well grounded in law documents. Their approach improved how the models behave, especially when the evidence is noisy or missing. They also found that training on legal data from one place helps the models do better in different legal systems too.
Open → 2609.14739v1

Group-relative reinforcement learning filtering can create phantom advantages

The Filter Metric is Safety-Critical: Phantom Advantages in Group-Relative RL under Shaped Rewards

Abstract: Group-relative policy optimization (GRPO and descendants) can discard no-contrast rollout groups through dynamic sampling, while practical implementations expose a configurable filter metric. We identify and quantify a metric-predicate mismatch under composite shaped rewards. When filtering follows the shaped training score rather than the task outcome, all-fail groups retain nonzero within-group spread and pass the predicate; standard-deviation normalization then promotes shaping differences among failures to full-size phantom advantages. In a controlled GSM8K comparison (Qwen2.5-1.5B, LoRA), no filtering and shaped-score filtering end at EM 0.080 +/- 0.112 and 0.040 +/- 0.008, whereas binary-outcome filtering holds 0.754 +/- 0.005 across four runs per arm (three default-seed reruns and one seed-123 run; mean +/- sample SD). On verl's native recipe/dapo trainer, holding model, data, reward and trainer fixed and changing only the metric, the score arm requires no batch refill in any of 40 observed steps and ends at EM 0.160; the accuracy arm refills in 29/40 steps and ends at 0.763. Both use the same custom shaped-reward hook and unmodified trainer/filter code. Prior work established shaping-induced amplification and all-fail filtering; our contribution isolates the metric-predicate semantic mismatch and directly instruments native deletion/refill telemetry. Across tested positive coefficients lambda in {0.1, 0.3, 0.5}, unsafe arms collapse; exploratory one-run cells reproduce the failure at 1.5B/7B on MATH and under GSPO, while disabling standard-deviation normalization avoids the observed collapse. Filtering under a composite reward should use a task-outcome signal whose semantics are independent of shaping.

Sat 12 SeptMachine Learning
The gist
Training AI models to make decisions often involves filtering groups of outcomes based on their scores. This paper finds that when filtering uses shaped reward scores (which mix task results with helpful hints) instead of just the actual task success, it can mistakenly treat failing groups as if some are better than others. This creates false advantages that hurt overall performance. The authors show that using the true task outcome as the filtering metric avoids this problem and leads to more reliable learning results.
Open → 2609.13866v1

ThinkPrior cuts wasted rollouts in reinforcement learning prompt selection

ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR

Abstract: In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term studied here depends on within-group reward variation. If all rollouts in a group are correct or all are wrong, their group-relative advantages are identically zero; these zero-advantage silent groups provide no reward-advantage gradient, yet uniform sampling spends 39% of a run's rollouts on them. History-based prompt selection must first spend target-policy rollouts to estimate difficulty, creating a cold start with rollout waste; ThinkPrior instead uses an external anchor in one offline pass to construct a zero-rollout difficulty prior before the first target-policy rollout. The verifier-scored anchor pass rate supplies an external-anchor initialization for a Beta posterior; ThinkPrior selects by expected learnability and then updates from training outcomes, changing neither the loss nor the optimizer. On Qwen2.5-Math-7B across sixteen seeds, ThinkPrior more than halves early silent groups and cuts wasted rollouts through step 30 by nearly a fifth, while we detect no difference in final accuracy. On this 250-prompt pool the fixed-budget result is a reallocation rather than a net saving. The measured ThinkPrior+DAPO composition reduces generated rollouts by 10.6% while both arms retain the same 3840-rollout update budget. The prior requires no target-policy rollout before the first selection, but the posterior thereafter uses target-policy outcomes.

Tue 8 SeptMachine LearningArtificial Intelligence
The gist
Selecting the best prompts to train AI models in reinforcement learning can waste many tries on unhelpful examples. The authors created ThinkPrior, a method that uses an offline 'anchor' test to guess how hard prompts are, reducing wasted attempts early on. This approach saves some training effort at the start without changing final accuracy. It reallocates effort more efficiently among prompts, leading to fewer wasted rollouts in early training.
Open → 2609.09075v1