Papers for

ai product development teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Adaptive distillation improves multi-turn reinforcement learning agents

TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL

Abstract: Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.

Mon 28 SeptMachine LearningArtificial Intelligence
The gist
Training AI agents to make good decisions over many steps is hard because they often get rewards only at the end, making early learning difficult. The authors note that previous methods mix teacher guidance with trial-and-error learning in a fixed way, which can hold the student back or cause poor learning. They propose TIDE, a method that adjusts how much the agent listens to a teacher versus learning from rewards both across the entire training process and at each step. By measuring when the student disagrees with the teacher, TIDE shifts more towards exploring better actions when it makes sense. Experiments show this approach helps train agents more effectively across tasks.
Open → 2609.35058v1

ZipBench cuts costs for testing large language models effectively

Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models

Abstract: Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluation unnecessarily expensive. Although recent benchmark compression methods (BCMs) can mitigate this cost, many strong BCMs rely on large collections of per-sample evaluation results from numerous LLMs to identify representative samples. Building such collections is also expensive unless they are already public, making these methods difficult to extend to newly released benchmarks. To address this challenge, we present ZipBench, a simple and low-cost BCM with theoretical error and rank-consistency guarantees. ZipBench evaluates only a small set of anchor LLMs, synthesizes pseudo evaluation results to broaden coverage, learns compact sample representations, and selects a small yet representative subset. Building on it, we create ZipBench Zoo, a collection of compact versions of 100+ benchmark proxies spanning text, multimodal, and agent tasks. These benchmark achieve mean absolute errors of 0.002--0.02 and average Spearman correlations of ~0.98 with the full benchmarks. Overall, ZipBench reduces the cost of both LLM evaluation and compact benchmark construction, lowering the barrier to broad LLM research for compute-constrained researchers. The code has been released in https://github.com/MilkThink-Lab/ZipBench.

Fri 11 SeptComputation and Language
The gist
Testing how good large language models (LLMs) are can take a lot of time and computing power because many benchmark tests are repetitive. The authors introduce ZipBench, a method that picks a small but smart selection of test examples by looking at just a few models and creating synthetic data to cover more cases. This approach keeps testing accurate but uses much less computing power and money. ZipBench makes it easier for groups with limited resources to evaluate and improve LLMs faster.
Open → 2609.12475v1