Papers for

robotic engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Efficient vision language control improves robot tasks with token caching

Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference

Abstract: Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.

Mon 28 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Robots that understand both pictures and words often need a lot of computing power to decide what to do. The authors found a way to make this faster without retraining the robot’s brain by saving and reusing certain important visual and text clues. Their method, called TVCache, picks the most helpful parts of the robot’s thought process to speed up decision-making, especially by focusing on how words and images work together. Tests showed robots using TVCache completed tasks more successfully and with less computing work.
Open → 2609.34319v1

Bioinspired estimator improves online tracking with partial sensor data

A bioinspired internal model-based online estimator for planar pursuit

Abstract: Bioinspired feedback controls for pursuit, tracking, and collective motion are often expressed in terms of the relative configuration between interacting agents. In practice, however, onboard sensors may not directly provide all quantities required for feedback control, necessitating estimation of unobserved quantities. This paper develops a bioinspired internal model-based estimator for reconstructing those quantities from partial sensory observations and known self-motion. State reconstruction is posed as an optimization problem that treats the relative kinematics as constraints and minimizes the disagreement between the internal model outputs and measurements from onboard sensors. Pontryagin's Maximum Principle is used to derive the necessary optimality conditions. A forward-backward algorithm is used to provide a numerical solution and a moving horizon formulation is employed for online implementation. The estimator is evaluated numerically against classical state estimators. Real-time implementation of the proposed framework on robotic hardware is demonstrated through two pursuit strategies.

Mon 21 SeptRobotics
The gist
Tracking and following moving targets often needs knowing relationships between agents, but sensors don’t always give all necessary information directly. The authors developed a way to estimate the missing information by combining a biological inspiration with a mathematical model that tries to best fit the available sensor data. They use advanced optimization techniques to solve this problem efficiently in real time. Their method works better than traditional state estimators and can be run live on robots to help them pursue targets more effectively.
Open → 2609.25470v1

Humanoid robot learns to swing across bars like primates

SwingBot: Learning Whole-Body Brachiation for Humanoid Robots

Abstract: Brachiation enables primates to move across overhead supports when ground paths are blocked, suggesting a complementary locomotion mode for robots operating in cluttered or hazardous environments. Bringing this capabil?ity to high-DoF humanoid robots is difficult because the controller must discover a long-horizon release-swing-capture sequence, coordinate alternating contacts with whole-body momentum, and act without reliable measurements of segment?relative displacement or hook-contact state. We present SwingBot, a learning framework for continuous humanoid brachiation with passive wrist hooks. Swing?Bot makes the task trainable by organizing learning around the structure of brachi?ation: biomimetic keyframes make rare release-swing-capture transitions reach?able during early exploration, and recurrent privileged-state estimation provides compact position and contact latents for deployment. Hardware experiments demonstrate continuous bar traversal and robustness to payload, external distur?bances and different bar spacings, showing that this formulation offers a practical route to whole-body robotic brachiation.

Wed 9 SeptRobotics
The gist
Moving through cluttered spaces can be hard for robots that walk on two legs. The authors developed SwingBot, a system that helps humanoid robots swing hand over hand, like monkeys on jungle gyms. Their approach breaks down this complex swinging motion into simpler steps and uses special learning to train the robot to keep its balance and grip. They tested SwingBot on real robots swinging through bars repeatedly, even with added weight and bumps, showing it can handle tricky movements.
Open → 2609.10283v1