Papers for

robotics control developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Learnable lens networks predict long-term dynamics with fewer parameters

LLN: Learnable Lens Networks for Parameter-Efficient Long-Horizon Dynamical Prediction

Abstract: Explicit residual connections of the form (x+f(x)), often combined with normalization layers, have become a standard strategy for training very deep neural networks. However, residual addition primarily provides an algebraic shortcut for gradient propagation, while leaving the evolution of feature geometry across layers largely unconstrained. We introduce Learnable Lens Networks (LLN), a physics-inspired architecture that replaces direct feature-space residual accumulation with learnable optical transport in an augmented position-angle phase space. Each layer alternates between free propagation, which provides an implicit transport path, and a learnable lens field that performs nonlinear trajectory transformation and focusing. Theoretically, we establish that LLN transport is globally invertible and volume-preserving for any differentiable lens field, with the implemented coordinate-wise Gaussian transport further satisfying symplecticity. Importantly, these structural constraints do not limit expressivity: with unrestricted embeddings and readouts, LLN retain universal approximation of continuous end-to-end maps. Experiments across diverse dynamical systems demonstrate that LLN improves long-horizon prediction while using substantially fewer parameters than same-depth comparators. Further analysis reveals stable depth-wise gradient transport and interpretable learned dynamics under the coupled propagation and refraction design.

Mon 28 SeptMachine Learning
The gist
Long-term prediction of how things change over time is hard for computers because errors build up. The authors introduce a new type of neural network called Learnable Lens Networks (LLN) that uses ideas from physics to better track these changes. Instead of simply adding updates directly, LLN simulates how light moves and bends, making predictions more stable and accurate over many steps. Their experiments show LLN can predict future states longer while using fewer resources.
Open → 2609.34493v1

Equivariant neural networks explained with graded polynomial theory

Graded Representation Theory of Equivariant Neural Networks

Abstract: Nonlinear activations can create equivariant interactions between irreducible representations that linear maps cannot. We use the Gaussian degree decomposition to extend ordinary polynomial degree to such nonlinear maps, and prove that for a fixed coordinatewise equivariant layer each degree factors into a polynomial determined by the linear maps and a scalar determined by the activation. This separates three distinct obstructions, coming from symmetry, coordinates, and activation.

Tue 22 SeptMachine Learning
The gist
Neural networks often use nonlinear activation functions, which help create complex interactions that simpler linear maps can’t capture. This paper explores how these nonlinear effects can be understood using a tool called Gaussian degree decomposition, extending the idea of polynomial degree to these more complex cases. The authors show that for a certain kind of neural network layer respecting symmetries, the nonlinear behavior can be separated into parts related to the network’s structure, the choice of coordinates, and the activation functions. This helps clarify what limits the network’s behavior based on symmetry and activation.
Open → 2609.25776v1

Convergence bounds clarify error sources in deep reinforcement learning

A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

Abstract: We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value function to targets from executed transitions and selects actions using a predictive model and the value function. For current observed-successor targets with fresh true-kernel outcomes, the conditional mean is $\mathcal{T}^βV$, which averages over behavior-policy actions. The Bellman optimality update is $\mathcal{T} V$. We decompose the update error into six residuals: fitting, transition reuse, target construction, replay, action selection, and exploration. Under $L^s$ concentrability, their $L^p$ norms ($p=s/(s-1)$) control expected $L^1$ policy loss. The bound explicitly weights residuals from only the last $H-1$ update blocks, plus an initialization term for shorter runs. We quantify the cost of a shared sampling distribution across horizon levels. For statistical error bounds of order $n^{-ν}$, we derive optimal continuous allocations and an integer allocation whose objective is within a factor $2^ν$ of the constrained optimum. A margin condition with exponent $α$ gives action error of order $Λ^{1+α/p}$, where $Λ$ combines network drift and score error; a one-step construction proves the exponent sharp. Bounds on the distance between frozen and optimal scores transfer an optimal-gap condition to frozen-iterate gap bounds while retaining the mass of optimal ties. Survival probabilities and coverage conditions at deployment yield bounds for policies selected with approximate scores. Separate spatial ReLU networks per horizon level give a conditional neural regression rate, and the finite-state case gives a log-free expected fit rate. These results give expected policy-loss consistency for the fixed-horizon generative-reset approximate-ERM procedure with exact action scores and provide an explicit residual-decay criterion for FIFO/interleaved SGD.

Wed 16 SeptMachine LearningRobotics
The gist
This paper studies how deep reinforcement learning algorithms that estimate future rewards can be mathematically guaranteed to get closer to the best possible decisions over time. The authors break down different sources of errors in the learning process and show how these errors combine and affect the overall performance. They provide formulas that explain how quickly and accurately the algorithm learns depending on factors like data reuse, action choices, and exploration. Their work helps understand when and why these algorithms work well and gives guidelines for ensuring consistent improvement. This kind of analysis supports building more reliable AI systems that learn from experience.
Open → 2609.18782v1

Autonomous driving improves by aligning decisions with future scene views

RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning. However, these World-Modeling VLAs rely on explicit future generation to learn such representations, thereby introducing two key limitations: additional training burden and inference latency. To address these limitations, we propose RAF-VLA (Representation Alignment with the Future), a VLA-based autonomous driving framework that shapes planning-relevant internal representations through direct guidance from future-frame representations. RAF-VLA employs Future-Aligned Supervised Fine-Tuning, in which a straightforward regularization aligns the policy's hidden states with future-frame representations obtained from a pretrained world encoder while learning driving actions. This simple alignment allows RAF-VLA to avoid the training burden and inference latency associated with future generation. Extensive experiments on the NAVSIM benchmark show that RAF-VLA achieves competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples seen. Moreover, RAF-VLA incurs only 3.8% training overhead and a negligible 1 ms inference overhead.

Tue 15 SeptRobotics
The gist
Autonomous cars need to plan their driving by understanding what will happen next on the road. Previous methods tried to predict future scenes explicitly, which slowed learning and made driving decisions slower. The authors present RAF-VLA, a way to align the internal thinking of the car with future road views without actually predicting those scenes. This approach helps the car learn faster and drive just as well, while being quicker in making decisions. Tests show it matches top driving systems but with less training and only a tiny increase in computing time.
Open → 2609.17728v1