Summary
Training AI agents to make good decisions over many steps is hard because they often get rewards only at the end, making early learning difficult. The authors note that previous methods mix teacher guidance with trial-and-error learning in a fixed way, which can hold the student back or cause poor learning. They propose TIDE, a method that adjusts how much the agent listens to a teacher versus learning from rewards both across the entire training process and at each step. By measuring when the student disagrees with the teacher, TIDE shifts more towards exploring better actions when it makes sense. Experiments show this approach helps train agents more effectively across tasks.
What this means in practice
- •For ai product development teams: Improve training of AI agents that interact repeatedly with users by adaptively combining teacher guidance and reward feedback for better long-term decisions.
- •For robotics engineers: Train robots more efficiently to perform multi-step tasks by dynamically balancing imitation of expert behavior and reward-driven learning during training.
Abstract
Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.