Adjoint meanflow matching speeds offline reinforcement learning actions

QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning

Machine Learning

Summary

Decision-making programs in AI often need to sample actions many times, which slows them down. The authors introduce QAMM, a technique that uses a special signal from a value estimator (critic) to train policies that predict average actions directly, requiring fewer steps to decide. This makes AI models faster while keeping good decision quality, tested on robot navigation tasks. Their approach combines ideas about how value changes guide actions with learning motion over short time intervals.

What this means in practice

  • For robotics engineers: Create robot navigation controllers that compute decisions quickly with fewer neural network calls by applying QAMM offline training.
  • For game ai developers: Develop game characters that choose complex actions in fewer computational steps using QAMM policies trained on past data.

Authors

Yuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou, Jiashu Hou, Ye Shi, Yanwei Fu

Abstract

Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.