Papers for

autonomous robotics teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Reinforcement learning improves by filtering bad training paths

Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning

Abstract: We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.

Mon 28 SeptMachine Learning
The gist
Training AI to make decisions often involves learning from examples called trajectories, but some of these examples can actually hurt learning. The authors show that existing methods to value data don't work well for reinforcement learning because the data is generated on the fly and there's no straightforward way to tell which examples are good or bad. They introduce a new method, Dynamic Trajectory Valuation (DTV), which uses the training process itself to figure out which training paths to keep or ignore, helping the AI learn better and more efficiently. Tests with popular reinforcement learning algorithms show that this approach makes training more stable and improves performance.
Open → 2609.35072v1

Posterior sampling achieves optimal exploration rates in reinforcement learning

Minimax-Optimality of Posterior Sampling for Reinforcement Learning

Abstract: Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayesian regret under arbitrary correlated priors. The difficulty is that a posterior-sampled transition model is coupled with its own continuation value. We overcome this with a common empirical transition reference that isolates the resulting value mismatch and a Bellman-based variance argument that controls it without an extra leading-order state-space factor. For finite-horizon, time-inhomogeneous tabular MDPs with unknown stochastic rewards, this yields the minimax $\widetilde{O}(\sqrt{SAH^3K})$ regret rate under arbitrary joint priors over rewards and transitions. The same proof principle gives the minimax $\widetilde{O}(d\sqrt{H^3K})$ rate for linear-mixture MDPs under arbitrary joint parameter priors.

Sun 27 SeptMachine Learning
The gist
Reinforcement learning helps computers learn to make decisions through trial and error. One popular method, posterior sampling, was simple but it wasn’t clear if it worked as well as possible in the most general situations. The authors show that this method actually does achieve the best possible performance in terms of learning speed and decision quality, even with complex prior knowledge. They used mathematical tools to carefully analyze why it works so well, confirming its effectiveness without extra assumptions.
Open → 2609.33246v1

Robot maps broken down into scenes to find changes fast

Online Geometric Change Detection via Scene Decomposition

Abstract: Autonomous robots are increasingly deployed on long duration single- and multi-session missions in dynamic environments, where the ability to identify environmental changes such as fallen trees or opened doors provides important contextual information for online planning. We propose a framework called Change Detection via Scene Decomposition (CDSD) for accurate online geometric change detection using LiDAR or RGB-D sensors. Recent advances in geometric SLAM have made it possible to generate dense, tightly aligned maps without post processing, but comparing global maps across entire sessions is computationally expensive and does not allow for single-session online change detection. CDSD instead spatially decomposes mapped environments into unique scenes where changes can be found efficiently by comparing dense, local subsets of the global map called submaps. As the first submap-based approach for geometric change detection, we identify and address the following core challenges: 1) identifying appropriate scenes for change detection that require minimal redundant information; 2) generating dense and representative submaps for each scene; 3) detecting changes between submaps with differing fields of view; and 4) processing detected changes for real-time map reconstruction. Results demonstrate our algorithm on custom datasets collected at the Army Research Laboratory facility in Graces Quarters, Maryland, and on open-source multi-session change detection datasets.

Tue 15 SeptRobotics
The gist
Finding changes in environments helps robots understand what’s around them and plan better as things change. The authors present a way to break big robot maps into smaller scenes, called submaps, to quickly spot differences like doors opening or trees falling. This method uses detailed 3D data from sensors and works while the robot is moving, instead of waiting until after a long mapping session. It’s tested on real and open datasets, showing it can detect changes efficiently and update the map in real time.
Open → 2609.17302v1