Papers for

robotics engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Unified model improves robot actions by linking memory prediction and execution

UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling

Abstract: Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) Transition ambiguity. Visually similar current observations may correspond to different manipulation phases and imply different subsequent transitions. (ii) Prediction--execution mismatch. A visually plausible predicted future observation does not necessarily correspond to a physically realizable transition. (iii) Experience--realization mismatch. A historically executable action pattern may not necessarily realize the intended transition in the current scene and therefore requires context-aware adaptation. Accordingly, we propose UniMPA, a Unified Memory-Prediction-Action model that addresses these problems through a shared action-grounded transition interface. (i) UniMPA introduces Persistent-Selective Future Prediction to resolve transition ambiguity by modeling the intended future state evolution. A persistent latent stream continuously tracks task-level progress, while a transition-critical pixel stream selectively resolves fine-grained interaction changes through memory-grounded prediction. (ii) To assess the physical executability of the anticipated transition, the predicted transition queries a temporal Visual-Action Memory Bank. The bank retrieves historically realized visual-action experience, grounding future prediction in executable evidence. (iii) To adapt executable experience to the current scene, an Action-Visual Memory Bank retrieves visually grounded action prototypes from historical action evolution. Prototype-Biased Flow then shifts the flow source toward a historically supported action manifold for context-aware refinement.

Thu 10 SeptRobotics
The gist
Robots often struggle to decide what action to take when what they see can mean different things, or when their planned actions don’t actually work in the real world. The authors introduce a new model called UniMPA that helps robots keep track of what they’re trying to do, predict what will happen next, and check past actions to make sure their plans can really be done. By combining memory of past experiences with predictions and current observations, the robot can better adapt its actions to the current situation. This approach aims to make robot manipulation more reliable and context-aware.
Open 2609.11875v1

Reinforcement learning plans near-optimally despite hard multi-step lookahead

Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead

Abstract: We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of $\ell$ actions before deciding its course of action. Although look-ahead can substantially improve achievable performance, it is known that optimal planning with multi-step transition look-ahead is NP-hard, but this hardness was established using discount factors arbitrarily close to one. It was therefore unknown whether the problem remains hard for any discount factor, and whether near-optimal planning can nevertheless be performed efficiently. We resolve both questions. First, we show that for every fixed rational discount factor ($γ\in(0,1)$), exact planning remains NP-hard. Second, we introduce a randomized polynomial-time approximation scheme for every fixed look-ahead depth. We then extend our approach to unknown transitions and stochastic rewards using optimism and variance-adaptive confidence bounds. The resulting algorithm achieves cumulative regret whose leading term matches classical tabular discounted RL up to logarithmic factors. Thus, although exact planning with transition look-ahead is NP-hard, efficient near-optimal planning and learning remain possible.

Thu 10 SeptMachine Learning
The gist
Deciding the best actions in a smart system that can see multiple future steps can be very difficult to solve exactly. The authors show that finding the perfect solution remains hard even when the system values the future less strongly. However, they also provide a way to find near-perfect solutions efficiently for any fixed lookahead depth. They extend this approach to learning when the environment is unknown, achieving performance close to the best possible. This means that even if exact planning is very complex, practical near-optimal planning and learning can be done efficiently.
Open 2609.11807v1

Robot hands learn to write in the air with a pen fast

Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian Estimation

Abstract: Dexterous in-hand manipulation of a grasped object with an anthropomorphic hand is an unsolved frontier for robot dexterity. The contact-richness and highly dynamic nature of object-hand interactions tend to require extensive modeling or data-collection efforts for learning-based approaches. Modern simulators used for reinforcement learning (RL) cannot fully replicate the required contact complexity, while collecting dexterous demonstrations for imitation learning (IL) remains an open problem. In this research, we present an embodied control approach based on real-time task Jacobian estimation of the combined hand and object system on the physical robot. Using only the CPU on a laptop, the proposed controller begins in-hand pen writing after approximately 18 s of initialization and continues to adapt online, without an analytic hand--object kinematic/contact model, simulation training, or precollected task demonstrations. We demonstrate that the same estimator/controller formulation works on three anthropomorphic robotic hand systems (one physical, two simulated) to show human-like, in-hand articulation of a grasped pen by an embodiment-independent formulation. Sub-millimeter in-plane precision (mean 0.6 mm across runs) is achieved across letters and shapes written in the air and on paper on a physical robot. To our knowledge, this is the first demonstration of an anthropomorphic hand writing arbitrary single-stroke trajectories with a grasped pen through purely in-hand motion, and it showcases an alternative to compute- and data-heavy approaches such as RL and IL for achieving dexterous manipulation through computationally simple and data-efficient algorithms.

Thu 10 SeptRobotics
The gist
It is very hard for robot hands to skillfully move and write with a pen inside their grasp because the way the hand and pen touch and move is complex and hard to model. The authors found a way for a robot hand to quickly learn to write letters by estimating how the pen moves based on the hand’s real-time motion, without needing lots of training or simulations. This method works on different robot hands and produces writing with very high precision. It shows a new way to get robots to do delicate tasks like writing using fast and simple calculations.
Open 2609.11775v1

Exoskeleton shared by human and robot improves dexterous hand learning

SEED-UMI: Sharing the Exoskeleton between human and robot for onE-to-one Dexterous demonstration

Abstract: Imitation learning for dexterous hands is bottlenecked by the difficulty of collecting contact-rich demonstrations that transfer faithfully to the robot. Prior wearable-exoskeleton systems record only on the human side and retarget via open-loop mappings calibrated in free space, which degrade under contact. We present SEED-UMI, a framework in which both the human and the robot wear the same exoskeleton: joint encoders become a physically shared measurement, and wrist cameras mounted to the exoskeleton observe the same outer mechanism during both human data collection and robot policy rollouts. This turns retargeting into paired cross-embodiment supervision and lets policies train directly on raw exoskeleton-centric wrist images, without segmentation or inpainting. On five contact-rich tasks, SEED-UMI achieves 3.0x greater data collection efficiency than exoskeleton-based teleoperation and a 70.0% average rollout success rate.

Thu 10 SeptRobotics
The gist
Teaching robots to use dexterous hands is hard because it's difficult to collect detailed demonstrations involving touch that robots can copy accurately. The authors created SEED-UMI, where both a human and a robot wear the same exoskeleton glove, sharing joint measurements and wrist camera views. This setup helps the robot learn more directly and effectively from the human movements, improving learning efficiency and success in handling tasks that require careful touch.
Open 2609.11753v1

Reflex informed learning improves muscle driven human walking control

Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion

Abstract: Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuromuscular control mechanism, while the reinforcement learning policy produces four biomechanically meaningful residual parameters to modulate key reflex gains and thresholds associated with hip swing, knee support, and ankle propulsion according to the current state. Experimental results demonstrate that the proposed framework generates physiologically plausible locomotion with improved kinematic accuracy and dynamic consistency, as well as better bilateral symmetry and stride-to-stride consistency under nominal walking conditions. The learned policy remains robust under muscle weakness and external perturbations without retraining.

Thu 10 SeptRoboticsGraphicsMachine Learning
The gist
Controlling human-like walking using muscle models is hard because it requires movements to be both realistic and adaptable to changes like muscle weakness or bumps. The authors combine a basic reflex system with reinforcement learning to adjust key muscle-related controls dynamically. This approach helps produce walking patterns that look more natural, stay balanced between legs, and handle disturbances without retraining. The resulting control method mimics how humans might adjust their muscle reflexes when walking under different conditions.
Open 2609.11733v1

ActSafeGuard keeps robot actions safe without lowering success

ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies

Abstract: Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSafeGuard, a differentiable and training-aligned safeguard layer for flow-matching based policies. ActSafeGuard integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component. Through an analytical ray-scaling operator design, ActSafeGuard enables boundary-aware gradients to guide the model to naturally learn constrained manifolds. Extensive experiments on multiple standard foundation backbones ($π_{0.5}$ and Fast-WAM) across various tasks demonstrate that ActSafeGuard consistently achieves a $100\%$ step safety rate while fully preserving or even boosting task success rates, providing a scalable and minimally invasive solution for safe embodied AI deployment.

Thu 10 SeptRoboticsArtificial Intelligence
The gist
Robots that interact with the real world need to make sure their actions don’t break physical limits, or else they risk causing harm or failing tasks. The paper’s authors created ActSafeGuard, a new method that helps robot action models learn to stay within these safety limits during training, not just at use time. This approach uses math to smoothly guide the robot’s decisions, ensuring all actions are safe while still succeeding at their goals. Tests showed the method reliably produced safe actions and sometimes even improved task success.
Open 2609.11697v1

Gait recognition improved with multimodal data and unified identity encoding

MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities

Abstract: Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmarks. We present MMGait, a large-scale multi-sensor benchmark that brings visible, infrared, depth, LiDAR, and radar observations into sequence-level correspondence. It provides diverse modalities spanning appearance, contours, geometry, motion, and body structure. Under a shared impostor-augmented protocol, we evaluate single-modal recognition, cross-modal recognition via directed retrieval, and multi-modal recognition using task-specific experts. Across settings, modality rankings vary with probe conditions, cross-modal alignment remains difficult, and fusion often provides complementary gains. This analysis exposes a scalability problem: individual modalities, modality pairs, and fusion configurations are typically handled by separately trained experts. We formulate Omni-Modal Gait Recognition, which unifies single-modal, cross-modal, and multi-modal recognition within a shared identity space. OmniGait++ uses modality-specific front ends followed by a shared identity encoder to preserve modality-dependent cues while learning comparable identity descriptors. An anchor-guided fusion module aggregates modality subsets of varying size without frame-level synchronization. A jointly trained checkpoint covers all three recognition settings and accommodates modality subsets of different compositions and cardinalities. Experiments show OmniGait++ remains competitive with task-specific experts in many shared settings and extends to higher-cardinality fusion unavailable to fixed-pair models. The results establish MMGait as a common testbed for heterogeneous gait sensing and demonstrate the feasibility of unified recognition under varying modality availability.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Recognizing people by the way they walk, called gait recognition, usually uses video images or simplified outlines. The authors point out that walking generates many different types of data like infrared, depth, radar, and more, which are rarely studied together. They created a big dataset called MMGait that combines all these types of data to compare and test gait recognition across different sensors. They also developed a system called OmniGait++ that can learn from and combine multiple sensor types at once, working well whether it uses one type of data or many. Their work shows it’s possible to have one system that adapts to different sensor combinations and still identify people accurately.
Open 2609.11601v1

Memory grounded planning improves real robot manipulation tasks

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

Abstract: Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.

Thu 10 SeptRobotics
The gist
Robots often struggle with tasks that require remembering lots of past details, because their usual methods only look at what’s happening right now. The authors present MaP-WAM, a way for robots to store past experiences as short plans, which helps them remember important details without getting overwhelmed. This new approach lets robots plan better and act more precisely over longer tasks, making memory use more efficient and improving success rates both in tests and on real robots. Their system keeps the amount of information the robot uses during actions fixed, so it stays fast even as tasks get longer.
Open 2609.11561v1

Humanoid walking improves with adaptive sensor noise handling

CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

Abstract: Humanoid locomotion across complex terrain demands forward-looking exteroception to anticipate obstacles, yet this signal is unreliable in real-world deployment, failing partially and intermittently. Existing perceptive policies often assume that depth observations remain clean and in-distribution, while recent attempts to unify perceptive and blind control typically route or switch between separate sub-policies, leaving recoverable information in partially corrupted depth unexploited. We instead propose CAP, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information. A coupled training recipe pairs a depth-noise curriculum on the world-model input with world-model feature dropout on the policy-facing latent, exposing the policy to failures across the entire perception-quality spectrum. In simulation, CAP matches or improves upon perceptive baselines when depth remains informative, and degrades more smoothly than a binary-switching baseline as perception worsens. On the Unitree G1, controlled trials and indoor-outdoor deployments demonstrate perception-robust locomotion under intermittent occlusion, real-sensor corruption, and outdoor depth artifacts.

Thu 10 SeptRobotics
The gist
Walking robots use special cameras to see obstacles ahead, but these cameras sometimes fail or give bad information. The authors created a method called CAP that can clean up noisy camera data and combine it with the robot’s own body sensors, so the robot can keep walking smoothly even when vision is spotty. They trained this system to handle various levels of camera noise, allowing the robot to adapt continuously rather than switching between on and off modes. Tests in simulation and on a real robot showed the approach helps the robot walk better when the camera data is partly broken or missing.
Open 2609.11553v1

Conditional transport boosts 3d deformable point cloud matching accuracy

BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration

Abstract: Reliable non-rigid point cloud correspondences are important for deformable anatomical registration, embodied perception and manipulation, and dynamic 3D reconstruction. Coarse-to-fine methods reduce computational cost by selecting the top-\(K\) coarse regions. However, this pruning may remove weak but correct hypotheses and restrict fine matching to an incomplete search space. We present \paper, a two-stage generative solver that maintains the complete soft matching matrix at both coarse and high resolutions. Stage~I uses denoising diffusion to estimate a global matching matrix in the compact coarse-resolution space. We then lift this matrix to high resolution while preserving its hierarchy. The lifted matrix is rank-bounded and block-constant. Stage~II refines it through a conditional transport bridge. We implement the bridge with two types of dynamics: a deterministic endpoint-parameterized conditional Flow Matching (CFM) ODE and a stochastic Brownian-bridge SDE inspired by Schrödinger bridges. Both variants share the lifted source, a time-conditioned transformer, and a matching-matrix endpoint predictor. Experiments on 4DMatch and 4DLoMatch show that both variants produce more accurate correspondences than the compared methods and improve downstream registration, with larger gains in low-overlap cases. They also improve cross-dataset generalization on CAPE and DeepDeform without target-domain adaptation while using the same deformation solver.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Matching points between two 3D shapes that deform is tricky, especially when the shapes only partially overlap. The authors introduce BridgeMatch, a new two-step method that first finds a rough global match and then refines it carefully without losing any possible correct matches. This approach uses advanced math tools called diffusion models and transport bridges to better guess which points correspond, even in challenging cases. Their experiments show it works better than previous methods and adapts well across different datasets.
Open 2609.11472v1

Frozen robot world models predict failures reliably and quickly

FARM: Reading Failure Signals from the Internal Predictive States of a Frozen Robotic World Model

Abstract: Reliable robot deployment requires online failure monitoring, yet existing monitors mainly derive risk from proxy signals or train dedicated monitoring components. We ask whether the internal predictive states of a frozen pretrained robotic world model already contain directly decodable failure information. Failure-Aware Readout from World Models (FARM) trains only a 33,985-parameter supervised readout over frozen VLA-JEPA predictive states, producing step-wise failure scores and causal trajectory risk. Five-fold out-of-fold evaluation across seven source tasks reaches 85.68/88.59 pooled AUROC/AUPRC, and FARM gives the best Seen performance among 15 matched baselines on the 10-task benchmark. Across four real-robot populations on PIPER X, SO-101, and Franka, fixed-readout transfer and readout-only adaptation test deployment shifts without updating the predictive backbone. FARM also discriminates failures from partial causal histories and adds 0.2256 ms mean CUDA latency once the frozen state is available. These results support frozen predictive world-model states as reusable features for causal, transferable, and low-overhead execution monitoring.

Thu 10 SeptRobotics
The gist
Robots need to quickly notice when something is going wrong to avoid bigger problems. This work shows that you can read signs of failure directly from a robot’s existing internal model that predicts the world, without changing that model. The authors designed a small add-on that looks at these internal predictions to score the risk of failure at each step, helping robots know when they might fail. Their method works well across many tasks and even transfers to real robots without needing to retrain the main model.
Open 2609.11445v1

Modular method simplifies closed-chain robot motion calculations

Modular Kinematic Reduction of Closed-Chain Mechanisms Using Path Assembly and Defect Homotopy

Abstract: Closed kinematic chains complicate modular modeling by coupling active and passive coordinates through nonlinear closure constraints. This paper presents a Path-Assembled Closure Differential Mapping (PACDM) framework for modular closure resolution and kinematic reduction. Each closure element compares two ordered transformation paths with common endpoints, with their mismatch expressed through the logarithm on SE(3) and the corresponding Jacobian assembled from local transformation derivatives. Multi-path modules are constructed from a minimal set of pairwise closure elements, while rank-revealing analysis selects locally independent scalar constraints. A defect homotopy recovers closure-consistent passive coordinates from approximate estimates along a feasible and regular continuation path. At regular configurations, implicit differentiation yields the local active-to-passive differential mapping, which is subsequently used in a predictor-corrector continuation procedure for prescribed motion. The framework is evaluated on a seven-degree-of-freedom heavy-duty manipulator containing two-path and three-path closed-chain modules. Comparison with Simscape Multibody yields trajectory root-mean-square errors below 8.5 x 10^-10 rad, while predictor-corrector continuation is approximately 45.8 times faster than applying defect homotopy at every trajectory sample.

Thu 10 SeptRobotics
The gist
Robots with closed loops in their arms are tricky to model because some parts move together in complicated ways. The authors created a new method that breaks down these loops into simpler pieces, making calculations easier and faster. Their approach uses math to measure differences along paths and fixes errors smoothly. Tests showed their method is very accurate and much quicker than older techniques. This could help design and control complex robotic arms more efficiently.
Open 2609.11338v1

Agent side memory guides robot actions for better long tasks

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

Abstract: Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrated geometry. These systems confound attribution: gains may come from richer observations or alternative motor tools, while failures may stem from either the policy or an under-specified language interface. We isolate this question through a deliberately constrained design: less tool breadth, but greater interface bandwidth. 2AM makes a multimodal Agent the sole holder of task memory and a single RGB-based, episodically stateless Action Model the sole executor of task-relevant motion. The Agent compiles interaction history into subtask language and optional 2D grasp, place, and move hints that bind its physical intention at different time scales. To teach this steerability to the VLA, we augment demonstrations with structured hint labels and train under condition dropout, spatial noise, and temporal jitter to tolerate imperfect Agent outputs. On LIBERO-Mem, without depth, online geometry, or planner-based object motion, 2AM reaches 76.3% average completion, a 61.5-point improvement over the strongest reported baseline of 14.8%, together with 63.0% relaxed and 11.8% strict success. These results show that task memory can remain Agent-side. They further show that Action Model capability depends not only on what the policy has learned, but on how precisely the Agent can steer it.

Thu 10 SeptRoboticsArtificial Intelligence
The gist
Robots trying to perform complex tasks over a long time need memory to remember what they have done. This paper shows it is possible to keep the memory in a separate part called the Agent, while the part controlling the robot's movements simply follows this guidance. The authors built a system called 2AM that uses only simple RGB cameras and no extra geometry or depth sensors. Their approach significantly improved success on a robot task benchmark compared to previous methods. This means robots can handle long, complicated tasks better by having a separate memory system that guides their actions.
Open 2609.11308v1

Dual latent space control improves robot policy learning efficiency

Beyond Noise Steering: Dual-Latent Space Reinforcement Learning for Generative Robot Policy

Abstract: Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generation process, resulting in performance degradation and inefficiency. To address this limitation, we propose a novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator. Specifically, our actor network predicts two distinct latent variables: an initial-noise latent variable that steers behavior generation, and an action-representation latent variable for intermediate feature modulation. Moreover, this representation latent variable is mapped to adapter features and ingeniously injected into the hidden states of intermediate action tokens via residual connections. Our dual-control design enables direct adjustment of action representations without updating the base policy. Experiments across generative policy architectures and robotic manipulation tasks show that DLSRL effectively accelerates online robot policy adaptation and achieves competitive performance. Our code is available at \href{https://github.com/xianchaoxiu/DLSRL}{https://github.com/xianchaoxiu/DLSRL}.

Thu 10 SeptRobotics
The gist
Robots that learn how to act by watching examples often use 'noisy' guesses that are hard to adjust precisely, which can slow down learning and reduce effectiveness. The authors propose a new method that lets the robot control both the starting guess and the details of the action as it develops. This double control helps the robot learn better and faster without changing its basic strategy. Their tests show this approach works well across different robots and tasks.
Open 2609.11270v1

Camera pose helps improve robot depth estimation from video

RIDE: Relocalization-Informed Depth Estimation with 3D Gaussian Splatting

Abstract: Render--match--PnP relocalization establishes correspondences between query image pixels and 3D map points for camera pose recovery, but their potential to support dense depth estimation is often overlooked. To exploit this geometric information, we present RIDE, which estimates dense metric depth from a robot's RGB stream. Given a metrically scaled 3D Gaussian Splatting (3DGS) model, RIDE combines sparse metric depth observations derived from PnP-RANSAC inlier correspondences with the geometric prior of a pretrained video-depth model. To handle uneven and intermittent observations, it integrates global and local depth correction with temporal memory, supporting depth estimation through short observation gaps after metric scale initialization. Trained on public RGB-D videos, RIDE is evaluated on robot sequences without fine tuning. Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.

Thu 10 SeptRobotics
The gist
Estimating how far away things are in a video is important for robots but can be tricky. The authors created a method called RIDE that uses the robot’s known camera position to help measure distances more accurately in videos. It combines information about where the camera is with clues from a trained model to guess depth, even when data is incomplete or noisy. This makes the robot’s depth perception more precise and consistent over time.
Open 2609.11079v1

Freehand sketching simplifies programming of robot swarm shapes

Freehand Sketching for End-User Programming of Robot Swarms

Abstract: Robot swarms are increasingly used in applications where accessible interaction with non-expert users is desirable. This paper investigates freehand sketching as an end-user programming interface for specifying robot swarm geometries. Users communicate spatial intent through a drawing, while the swarm autonomously extracts target formation points, constructs a rigid formation graph, assigns robots to formation nodes, and executes distributed formation control with a guarantee against unintended reflected formations. The resulting sketch-to-swarm framework is evaluated through a human study examining the usability of freehand formation specification. Twenty participants generated $42$ geometric shapes, and the interface achieved a mean System Usability Scale score of $84.25$, which conventionally indicates high perceived usability. The results support freehand sketching as an intuitive interaction abstraction for human-swarm collaboration without requiring robotics or programming expertise.

Thu 10 SeptRobotics
The gist
Drawing is a natural way for people to show how they want groups of robots to arrange themselves. The authors developed a system where users sketch a shape by hand, and the robot swarm figures out where each robot should go to match the drawing. This system avoids confusing robot formations by making sure the shape is not mistakenly flipped. In tests, people found the sketching system easy and natural to use, even without robot or programming knowledge.
Open 2609.11078v1

3D printing hinges that sense multiple movements without extra parts

X-Hinges: 3D Printing Self-Sensing Compliant Mechanisms for Continuous and Multi-DOF Motion Sensing

Abstract: We present X-Hinges, a design and fabrication method for self-sensing compliant mechanisms based on multi-material FDM 3D printing. By co-printing two conductive filaments of different conductivities within a compliant body, we embed resistive sensing elements directly during fabrication without post-assembly, enabling continuous motion sensing across multiple degrees of freedom in a single print. The structure supports three degrees of freedom, each equipped with a dedicated sensing element configuration for multi-DOF motion estimation. We develop a precision data acquisition system and data-driven regression models that enable continuous, real-time motion sensing. We also introduce an interactive design tool for customizing the geometry, mechanical properties, degrees of freedom, and sensing configurations of X-Hinges. The tool also supports augmenting existing 3D models with self-sensing structures, endowing ordinary objects with continuous multi-DOF sensing capabilities. Finally, we present a set of application examples demonstrating the capability of X-Hinges for fabricating personalized interactive interfaces.

Thu 10 SeptHuman-Computer Interaction
The gist
Devices often need sensors to know how they move, but these usually require extra parts and assembly. The authors created hinges that can be 3D printed with special materials to sense their own motion in several directions all at once. By mixing different conductive plastics during printing, their hinges measure movement continuously without adding sensors later. They also built tools to design and use these self-sensing hinges in various objects.
Open 2609.11077v1

Passive arm stiffness affects payload stability in quadruped walking

Gait-Dependent Effects on Quadruped Locomotion for Load-Carrying using Passive Mechanism

Abstract: Passive mechanical interfaces offer a lightweight alternative to actuated manipulators for quadruped payload carrying, but their impedance directly couples the payload dynamics with the locomotion pattern. This paper analyzes how passive-arm stiffness-damping selection affects payload-carrying locomotion under different gait and payload conditions. We compare damped and underdamped passive-arm impedance configurations in simulation during flat-ground locomotion. For crawl gaits, where the support polygon remains well defined, the results show that underdamped impedance increases passive-joint oscillations and can reduce the ZMP margin with respect to the support polygon. Trot is retained as a dynamic excitation case for the passive arm, but it is not used for direct ZMP-margin stability comparison. The results are summarized in gait-payload-stiffness-damping maps, where ZMP-margin reduction is evaluated for crawl gaits and trot is retained only as a passive-arm excitation case.

Thu 10 SeptRobotics
The gist
Carrying loads on four-legged robots using simple, spring-like arms is lighter than using powered arms, but the way these arms move depends on the walking style. The authors studied how adjusting the arm’s spring and damping settings changes carrying stability during different walking patterns. They found that with a slow and stable walk (crawl gait), making the arms less damped causes more swinging and can make the robot less stable. The study provides maps that help choose arm settings for different walks and payloads to keep the robot balanced.
Open 2609.11059v1

Vision language models struggle to decide when to gather more physical data

New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

Abstract: A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial experiment may provide enough information to answer, or the model may need another measurement, such as the object's mass, friction, restitution, or spring stiffness. We study whether vision language models can decide when to answer immediately and, when more evidence is needed, which experiment to perform. Current physical reasoning benchmarks usually evaluate only the final answer, so they do not directly measure this decision-making ability. We introduce a controlled evaluation where each problem provides one measurement image and four possible physical worlds created by combining two possible masses and two possible values of another relevant property. The model must either stop and answer or select the cheapest additional experiment that can resolve the question. We construct matched problem pairs where changing either the observed measurement or the question changes the optimal action. Since all possible worlds and experiment costs are known, we can explicitly determine the optimal choice. Across six open models and 144 physical parameter sets, direct responses repeat the same action for 95.1% to 100% of image pairs even when the correct action changes. Brief reasoning improves action switching, but the best model makes both decisions correctly for only 5.9% of image pairs. Additional analysis reveals failures in measurement interpretation, physical reasoning, and response formatting. By evaluating evidence selection separately from final answers, our benchmark reveals limitations in physical reasoning that conventional answer accuracy can overlook.

Thu 10 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language
The gist
Sometimes, machines look at a picture from an experiment to answer a question about how something moves or reacts. The authors tested if these machines can decide when they have enough information or when they need to ask for more experiments to find out. They found that current vision language models usually don't change their choice even when the right answer depends on new data. This shows that these models have trouble figuring out when to gather new evidence, not just giving the right answer.
Open 2609.11022v1

Clarifying what makes AI count as an agent and how to measure it

Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks

Abstract: The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence. For each dimension, we examine how the underlying capability has been conceptualized across prior work and synthesize the metrics, benchmarks, and evaluation frameworks used to assess it. This review provides a structured account of the current landscape of agent evaluation, highlighting both established approaches and areas where evaluation remains limited or inconsistent. We additionally introduce the Agent Compendium, a public-facing digital resource that organizes and extends the evaluation methods identified through this review. Together, the survey and compendium provide a common structure for evaluating and comparing agent capabilities across AI systems, supporting more reproducible research, clearer communication, and more systematic study of artificial agents.

Thu 10 SeptArtificial IntelligenceMultiagent Systems
The gist
It can be hard to say exactly what an AI agent is because different researchers use the word differently. The authors looked at five key features that AI agents might have, like how they interact with their world, learn, act on their own, aim for goals, and keep consistent over time. They gathered many ways these features have been measured before and made a resource that collects these measures together. This can help people compare and test AI agents more clearly and fairly.
Open 2609.11018v1

Topological stages improve long term robot control across embodiments

Topological Necessities: Mechanism-Invariant Strategic Subgoals for Cross-Embodiment Goal-Conditioned Control

Abstract: Long-horizon goal-conditioned reinforcement learning delegates control to a high-level module that proposes subgoals, but existing subgoals are implicit byproducts of value functions or latent actions, tied to the executor that produced them. We study a different object: a route-conditioned order of unavoidable stages that every successful executor must traverse, recoverable from offline trajectories and belonging to none of them. Its defining properties are topological: an unskippable stage is a separating set that every admissible path must cross, and a loop in free space forces a route choice. We read the two by homology in dimensions 0 and 1 over a transport-weighted carrier built from successful trajectories, yielding an enumerable gate set with shell-level certificates; the certified gates are what we call topological necessities. Certified gates enter the decision loop as a recursive topological gate hierarchy. Under a fixed, isomorphic free space, the object survives executor replacement: gates frozen on PointMaze data transfer without retraining to Ant and Humanoid, attaining the highest Humanoid aggregate under a unified interface (96.1), with +36.0 over a map-privileged reference on the multi-route task (p=1.4e-5); the planner saturates PointMaze (100+/-0) and matches or exceeds the strongest baselines on AntMaze (giant +22.9) and Kitchen (+15.8/+12.6).

Thu 10 SeptMachine LearningArtificial IntelligenceRobotics
The gist
Many robot control systems break big tasks into smaller steps called subgoals, but these steps are usually specific to the robot that learned them. This paper introduces a way to find essential passing points or stages that any robot must go through to complete a task, based on the shape of the environment and past successful attempts. These stages don’t depend on the robot itself and can be used across very different robots, like simple point agents or humanoid robots. The authors show that using these topological necessities improves performance and transfers well between different robot types.
Open 2609.11014v1

World model updates measured for utility using counterfactual deployments

Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation

Abstract: Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment had it held its parameters. We introduce the fork ledger, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers. It evaluates both continuations on the same episodes and records $ΔR = R_{\mathrm{update}} - R_{\mathrm{hold}}$. Always applying one fixed update mechanism lowers return on all three simulated control tasks: CartPole ($-144.0$; checkpoint-bootstrap $95\%$ CI $[-185.4,-116.1]$, against a converged return near $650$), Walker ($-82.8$; $[-101.1,-61.7]$) and Cheetah ($-18.6$; $[-29.0,-6.6]$). Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the $693$ of $720$ that did not collapse, CartPole and Walker are unchanged in sign ($-113.4$ and $-82.1$) and Cheetah becomes unresolved ($-3.9$; $[-17.5,+13.0]$). The task is the unit of inference: each contributes $240$ attempted forks over five pretrained checkpoints crossed with two drift directions. The ledger makes counterfactual utility observable for a fixed mechanism, allowing triggers to be judged by the updates they select rather than by surprise detection alone.

Thu 10 SeptMachine Learning
The gist
When a computer program that predicts how the world works gets new information, it needs to decide if it should update its understanding. The authors introduced a way to test the value of each update by comparing what happens if the model updates versus if it stays the same at the same moment. Their tests on simulated control tasks showed that always updating can sometimes make performance worse. This new method helps decide when updates really improve predictions rather than just reacting to surprises.
Open 2609.10954v1

Imle-vla speeds robot action with single-step vision language model

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

Abstract: Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 10 Euler steps in $π_{0.5}$. This creates an inference bottleneck that produces stop-and-go movement in the robot and slower task completion. We introduce IMLE-VLA, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). The cIMLE objective promotes multimodal action coverage, avoiding the mode collapse of naive regression heads while eliminating multi-step sampling entirely. When IMLE-VLA is applied to $π_{0.5}$, it increases inference frequency 3.67x (55 Hz vs. 15 Hz), enabling up to 11x higher action throughput. On the 40-task LIBERO benchmark, IMLE-VLA achieves the highest average success rate (98.0%) among all baselines while leading in inference frequency. Under the test-time perturbations of LIBERO-plus, IMLE-VLA retains $π_{0.5}$'s robustness while other baselines degrade sharply, confirming that the cIMLE head preserves generalization. Real-world experiments on a Franka Emika Panda across four tasks demonstrate smoother motion (2.2x to 3.0x lower jerk) and faster task completion, with IMLE-VLA outperforming $π_{0.5}$ on every task and reducing average VLA inference time per episode by 3.9x to 6.6x. Videos and code are available at https://kianhk6.github.io/IMLE-VLA/

Thu 10 SeptRoboticsComputer Vision and Pattern Recognition
The gist
Robots that understand images and language often take many small steps to decide what to do next, making them slow and jerky. The authors propose IMLE-VLA, a method that generates robot actions in one quick step instead of multiple steps. This makes the robot move smoother and finish tasks faster without losing accuracy or flexibility. They tested IMLE-VLA on many tasks in simulation and on a real robot, showing it is faster and just as reliable as previous methods.
Open 2609.10915v1

Robot motion planning improved with gradient-friendly inverse kinematics

Planning along Differentiable Charts of Constraint Manifolds with General-Purpose IK Solvers

Abstract: Planning trajectories for robot manipulators under kinematic equality constraints restricts feasible motions to a measure-zero submanifold of the configuration space, requiring special algorithmic treatment. A promising strategy is parametrizing the set of feasible configurations using analytic inverse kinematics (IK). Bespoke analytic IK functions can be written to be differentiable, a necessary property for gradient-based trajectory optimization. But the vast majority of IK functions are computed by automated meta-solvers like IKFast, and are difficult to modify for differentiability. We present a new approach for computing gradients of analytic IK parameterizations: we leverage the inverse function theorem to recover the desired gradients from the ordinary forward kinematic Jacobian. Furthermore, we present a least-squares domain extension and an optimization-amenable description of the reachability constraint, which preserves gradient signal outside the reachable workspace. We demonstrate the efficacy of our approach through numerical experiments and downstream tasks, including a hardware demonstration of an RB-Y1 picking up a box and placing it on a table. Project website: https://cohnt.github.io/inverse-function-theorem-parameterization/

Wed 9 SeptRobotics
The gist
Planning robot arm movements gets tricky when the arm must follow exact rules, because the possible positions form a tiny set in a huge space. The authors found a way to calculate how to adjust robot arm positions smoothly using math called gradients, even when using common inverse kinematics tools that weren't designed for that. They do this by cleverly using the forward movement formulas backwards, allowing robots to plan better and reach tricky positions more reliably. This method was tested in simulations and real robot tasks like picking up and placing boxes.
Open 2609.10905v1

Vision language models tested on coastal and underwater scenes

Evaluation of Vision-Language Models Across Diverse Coastal Environments

Abstract: Vision-language models (VLMs) enable robotic per- ception by associating visual observations with natural-language concepts. Yet their performance in coastal environments remains largely unexplored. We introduce a densely labeled coastal dataset containing more than 1,000 images collected across seven missions in three regions of Oahu, Hawaii, with 18 semantic classes and over 7,400 annotated instances. We evaluate seven modern VLMs through three complementary experiments mea- suring text-to-mask, mask-to-mask, and mask-to-text alignment. Broad landscape classes are generally recognized more accurately than conventional object and coastal classes, with coastal con- cepts presenting the greatest challenge. However, comparisons of shared conventional classes across coastal and terrestrial datasets reveal no consistent performance difference attributable solely to environmental context. Mask-to-mask matching also remains similar across conventional and coastal classes, while alternative textual labels substantially improve recognition of several coastal concepts. These results suggest that lower performance on coastal classes (at least on the objects/query categories evaluated) is heavily influenced by segmentation and linguistic representation.

Wed 9 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Vision-language models help robots understand pictures by linking images to words. This paper studies how well these models work with photos from coastal areas in Hawaii, which can be different from land scenes. The authors made a special dataset of over 1,000 coastal images with detailed labels and tested seven popular models. They found that models recognize broad landscapes better than specific coastal objects, and special labels can help improve recognizing some coastal things. The challenges come mainly from how objects are segmented in images and described by words.
Open 2609.10855v1

Robotic pianist achieves expressive performances like human players

Expressive Robotic Pianist: Mastering Complex Piano Repertoire with Graph-Mimic and Musical Dynamics

Abstract: Enabling robots to perform musical instruments with human-level expressivity represents a frontier in bridging the gap between mechanical precision and artistic interpretation. Despite advances in robotic dexterity, replicating the fluid finger transitions and nuanced dynamic control characteristic of human pianists remains a significant challenge. Through a reinforcement learning-based control framework, we demonstrate that a dexterous robotic hand can achieve high-fidelity performance across a diverse piano repertoire. Central to our approach is a graph-based optimization strategy that guides the robot to generate natural pre-press and key-press fingering strategies that closely resemble human movement patterns. To achieve expressive sound production, the control system is coupled with a physics-inspired acoustic model that modulates keypress velocity to accurately reproduce the dynamic variations specified in musical scores. Quantitative evaluations demonstrate that our expressive control model significantly outperforms baseline methods in both finger morphology similarity and dynamic velocity accuracy. In a perceptual test involving participants from diverse listener groups, performances generated by our system are significantly preferred over baseline robotic performances and are indistinguishable from human performances for non-professional audiences. Furthermore, extensive experiments across multiple musical styles confirm that our method maintains high note-level accuracy while achieving expressive performance. Our approach provides a robust pathway for robotic systems to move beyond mere mechanical accuracy, elevating robotic musicianship to a level of expressive performance comparable to human pianists.

Wed 9 SeptRobotics
The gist
Playing piano like a human is hard for robots because it requires smooth finger movements and control of how loud or soft the notes sound. The authors created a system where a robot hand learns to play piano pieces more naturally by following finger movements similar to humans. Their system also adjusts how hard it presses the keys to match the music’s dynamics, like loudness changes. Tests show people prefer the robot’s performances over earlier robot attempts, and non-experts sometimes can’t tell the difference from a human pianist.
Open 2609.10844v1

Graph connectivity boosts hierarchical reinforcement learning with dense rewards

From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs

Abstract: The integration of graphs with Goal-Conditioned Hierarchical Reinforcement Learning (GCHRL) has received increasing attention, as graphs naturally encode task hierarchies for effective subgoal sampling. However, existing methods often overlook intrinsic connectivity information, failing to fully leverage the underlying topology for efficient learning. Most graph-based GCHRL methods use the graph as a stochastic sampling tool rather than as an environmental model that encodes connectivity and state-accessibility information. This limitation is particularly acute in quasimetric environments, where the inherent asymmetry of state transitions poses a fundamental challenge to stable policy learning and robust path planning. In this paper, we address these problems by introducing a state connectivity model designed to predict pairwise state connectivity strength in asymmetric environments. We transform these connectivity strengths into scalar auxiliary dense rewards, providing continuous guidance across multiple hierarchical levels. We demonstrate that our proposed framework, Graph-Guided Quasimetric Dense Reward (G2QDR), can theoretically be integrated into any existing GCHRL architecture, and the state connectivity model is efficiently implemented via a neural network trained on a directed state graph generated during exploration. Empirical results across a wide range of sparse reward environments indicate that, in general, G2QDR can enhance the performance of baseline GCHRL approaches with acceptable computational overhead.

Wed 9 SeptMachine Learning
The gist
Many reinforcement learning systems learn tasks by setting and reaching smaller goals, but they often miss important connections between places or states that help guide learning. The authors offer a method that creates a map capturing how strongly different states are connected, even when moving between states isn’t always the same in both directions. They turn these connections into helpful, continuous rewards that guide the learning process more smoothly. Their approach works alongside many existing methods and shows better performance in tasks where rewards are initially rare.
Open 2609.10781v1

Optimal equation helps balance learning stability and flexibility

A Bellman Optimality Equation for Plasticity

Abstract: In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et al. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent's observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optimizing plasticity within Markov decision processes. We show that there exists a Bellman optimality equation for optimizing plasticity similar to previous work for empowerment.

Wed 9 SeptMachine Learning
The gist
Balancing learning stability and the ability to adapt, called plasticity, is a key challenge in AI systems that learn continually. The authors build on previous work that defined plasticity and empowerment in terms of information flow between an AI agent’s actions and observations. They show there is a mathematical equation, like the famous Bellman optimality equation, that can help optimize plasticity in decision-making processes. This is a first step toward better managing how much AI agents change their behavior over time.
Open 2609.10776v1

Generative radar depth estimation works well through smoke and fog

GRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation

Abstract: Dense 3D depth perception fails under smoke, fog, and darkness because optical sensors cannot penetrate airborne particulates. mmWave radar remains usable and measures range accurately under these conditions, but its small aperture limits angular resolution. We present GRADE, which grounds a pretrained generative prior in single-frame radar geometry to estimate high-fidelity metric depth. GRADE first maps raw 4D radar spectra to coarse metric depth. A latent diffusion backbone then recovers structural detail while conditioning every denoising step on this estimate. A pixel-space adapter uses residual camera cues when available and is trained across clear, smoke-degraded, and occluded inputs so the full output approaches the radar-conditioned path as visibility degrades. Trained and evaluated on ~95K frames across 12 buildings with real smoke, GRADE achieves an MAE of 0.303 m in clear scenes and 0.313 m under smoke, outperforming existing baselines. Code and datasets are available at https://phi-lab-rice.github.io/GRADE.

Wed 9 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Seeing through smoke, fog, or darkness is hard for regular cameras because light gets blocked or scattered. The authors show that radar, which uses radio waves, can measure distances well even in these conditions, but it usually captures coarse images. They developed GRADE, a method that uses a kind of smart image generation to turn low-resolution radar data into detailed 3D depth maps. This approach works well both in clear air and through real smoke, improving on other techniques.
Open 2609.10756v1

Robotized human videos improve robot learning with large scale data

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Abstract: Human video datasets have emerged as a compelling alternative to expensive real-robot data, offering rich diversity at scale. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can provide effective and scalable supervision for pretraining vision-language-action (VLA) policies. To this end, we develop a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. Using this pipeline, we construct the HuRo dataset, comprising about 630K robotized episodes and 142M processed frames from five human-video sources. Across four real-world manipulation tasks, increasing robotized pretraining scale improves overall completion from 51.5% to 80.3% and OOD completion under spatial and visual shifts from 34.9% to 72.2%. Ablations further show that visual robotization improves OOD robustness and that end-to-end pretraining with retargeted actions outperforms visual-only transfer. Code and data are released on our website: https://3587jjh.github.io/HuRo.

Wed 9 SeptRoboticsComputer Vision and Pattern RecognitionMachine Learning
The gist
Robots need lots of example videos with matching actions to learn how to do tasks, but collecting real robot videos is expensive. The authors developed a method to change many different human videos into videos that look like robot views and include robot actions. This large video dataset helps teach robots better, making them more likely to complete tasks in new scenes. Their work shows that using both the robot-like video and the matched actions together is better than just using video alone.
Open 2609.10706v1

Image to 3D models improve shape accuracy using partial test time data

Guiding Image-to-3D Generation with Test-Time Partial Observations

Abstract: Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the available observations, limiting their use in applications that require geometric fidelity. In many real-world settings, however, partial geometric observations of the object may be available at test time. We introduce a training-free framework for incorporating such evidence into pretrained image-to-3D generative models without retraining or finetuning. To do this, we guide generation using a ray-consistent observation likelihood defined over the model's occupancy representation, combining surface occupancy and free-space evidence. Applied to SAM 3D and its multi-view extension, our approach substantially improves geometric fidelity across different levels of observability, as well as visual quality. Our results demonstrate that pretrained image-to-3D models can effectively integrate partial geometric observations through explicit test-time guidance, complementing their learned generative priors without modifying the underlying model.

Wed 9 SeptComputer Vision and Pattern Recognition
The gist
Creating 3D models from a single 2D image is tricky because the shape details are often uncertain. The authors found a way to use extra 3D clues available during testing to make the models more accurate, without changing the original model. They do this by comparing what the model predicts with these partial 3D clues using a math approach that checks if rays hit the expected object or empty space. This makes the 3D results better in both shape accuracy and visual quality.
Open 2609.10531v1

Semigroup-jepa improves physics prediction and control in varied gravity

Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization

Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning and jointly training an encoder and predictor through an autoregressive latent rollout. To evaluate the model's ability to generalize out of distribution, we design dynamical tasks under different gravitational fields that, despite obeying the same physical law, exhibit qualitatively different dynamics, ranging from floating motion in weak gravitational fields to rapid bouncing in strong ones. In contrast to DINO-WM, SG-JEPA reduces open-loop prediction error by up to 2 times on two-dimensional datasets, and increases control success rate up to 2.5 times for three-dimensional robotic datasets, for which we train independent diffusion policies. To explain this advantage, we develop a linear feature model that separates local law-conditioned error from its recursive amplification under rollout. Guided by this model, we find that back-propagating the multi-step rollout loss into the representation trains the encoder to keep the features that the predictor can carry forward, and that those are the features the dynamics depend on, so most of the gain comes from the encoder learning better features rather than from the predictor learning better dynamics. See project page at https://sg-jepa.github.io.

Wed 9 SeptMachine LearningArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Predicting how objects move under different gravity levels is challenging for AI models. The authors created Semigroup-JEPA, a model that learns to represent and predict physics more accurately across a wide range of gravity environments. It does this by training the system to keep track of meaningful features that help forecast future motion, leading to better predictions and more successful robot control tasks. Their tests show this approach cuts prediction errors and improves robotic task success rates when gravity changes.
Open 2609.10464v1

Vision language action models improve robot motion understanding with frequencies

Frequency-Conditioned Flow Matching for Vision-Language-Action Models

Abstract: Robot actions are temporally correlated trajectories whose frequency components encode motion at different scales with highly non-uniform energy distributions. Yet Flow Matching--based vision-language-action (VLA) models typically generate actions in temporal coordinates, without explicitly modeling or systematically leveraging this frequency heterogeneity. We introduce \emph{FreqFM}, a frequency-conditioned Flow Matching framework for VLA models. It raises action frequency from an implicit trajectory property to an explicit conditioning dimension that spans the entire generation pipeline. Concretely, in DCT frequency coordinates, FreqFM constructs a spectrum-matched source distribution, adaptively balances the objective across frequencies, and constrains per-frequency guidance residuals using the corresponding reference transport scales. FreqFM integrates into existing Flow Matching action experts without changing the VLA backbone. Across LIBERO, LIBERO-Plus, and VLA-Arena, FreqFM consistently improves performance, including a 9.3-point gain on LIBERO-Plus, and further demonstrates its effectiveness on six real-robot tasks.

Wed 9 SeptRobotics
The gist
Robot actions happen over time and include moves at different speeds and sizes, which can be seen as having different 'frequencies.' Usually, computer models guess robot motions without paying much attention to these different frequencies. The authors created a new method called FreqFM that explicitly uses these frequency details to better predict actions. This approach fits into existing models and helps robots perform tasks more accurately in tests and real robots.
Open 2609.10405v1

Humanoid robots learn to walk on loose sand and gravel terrain

Learning Terrain-Adaptive Humanoid Locomotion on Granular Terrain

Abstract: Humanoid locomotion on granular terrain remains a significant challenge due to its complex foot-terrain interaction dynamics that are difficult to model. Existing approaches either ignore granular contact dynamics or incorporate simplified normal force models with heuristic tangential components. In this work, we present a physics-grounded granular contact model based on three-dimensional resistive force theory (3D RFT) and efficiently simulate granular terrain for reinforcement learning (RL) training. Unlike traditional rigid contact models and simplified granular contact models with ad-hoc heuristics, our contact solver produces physically accurate granular intrusion dynamics without resorting to heuristics. It captures realistic penetration and tangential drag during training, enabling the policy to learn behaviors that transfer reliably to real-world granular terrain where rigid contact models fail. To adapt to varying terrain conditions, we train a terrain-adaptive locomotion controller via teacher-student RL, using a variational autoencoder to encode terrain information into a compact latent representation. Simulation studies using material point method (MPM) with NVIDIA Newton demonstrate that our method generalizes to unseen granular terrains, achieves a significantly higher success rate than baselines, and demonstrates zero-shot terrain identification and adaptation. We further validate our approach through extensive hardware experiments across diverse real-world granular terrains including basalt, dry sand, and beach sand. To the best of our knowledge, this is the first demonstration of agile humanoid locomotion on real-world granular terrain. Project page: https://humanoid-gm-locomotion.github.io/HUMANOID-GM/

Wed 9 SeptRobotics
The gist
Walking on loose materials like sand and gravel is hard for robots because the ground changes shape when stepped on. The authors created a new computer model that better predicts how feet push into and move through these loose surfaces. They trained robots in simulations using this realistic model, allowing the robots to learn how to walk steadily on these tricky grounds. The robots then successfully walked on real sand and gravel without needing extra adjustments. This is the first time humanoid robots have done this well on real loose materials.
Open 2609.10286v1

Legged robots use frame coding to handle rough terrain better

Frame-Coded Legged Locomotion over Noisy Terrain

Abstract: Open-loop multilegged locomotion over rough terrain has been interpreted as matter transport over a noisy channel: leg-ground interactions are discrete basic active contacts, terrain deletes or perturbs those contacts, and spatial redundancy concentrates the resulting thrust and arrival time. That construction is repetition-like because every module carries the same scalar locomotion task. It consequently provides neither a positive task rate nor a decoder that changes with the surviving contact set. Here we formulate locomotion instead as a quantized finite-frame expansion with erasures. A d-dimensional body-level command is mapped into N>d heterogeneous local contact commands. Rough terrain erases or corrupts frame coefficients, while a contact-gated compliant morphology physically realizes the weighted active-subframe decoder. For a linear-Gaussian model, mechanical equilibrium is exactly the posterior mean, tangent stiffness is posterior precision, and mechanical compliance is posterior covariance. Equal-norm Parseval frames are shown to be minimax optimal against one missing contact, two-contact robustness is governed by frame coherence, and a harmonic frame gives a directly realizable gait family. For independently surviving contacts of probability q, random Gaussian gait frames admit exact reconstruction at every analog dimension rate R<q, with a binomial reliability exponent, whereas recovery of arbitrary commands is impossible for R>q. Residual contact noise yields an asymptotic per-mode amplification 1/(q-R) and a vanishing mechanical stiffness margin at the threshold. An information-locomotion inequality and an exact incremental-redundancy rule direct the next gait component toward the softest task-relevant unresolved mode. The resulting analog frame-coding theorem establishes a finite relative redundancy and converse as part of a fundamental limit theory of legged locomotion.

Wed 9 SeptInformation TheoryRobotics
The gist
Walking robots often struggle when the ground is uneven or slippery because their legs don’t always get a good grip. This work treats each leg’s contact with the ground like a small part of a code that can be lost or corrupted by rough terrain. The authors found a mathematical way to spread the robot’s movement commands across many leg contacts so that even if some legs slip or fail, the robot can still walk steadily. Their approach uses ideas from signal processing and statistics to understand and design how robots react to missing or noisy foot contacts.
Open 2609.10273v1

FolDeX provides real-robot benchmark for long deformable object manipulation

FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

Abstract: Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-object manipulation, where policies must track changing states and execute reliable multi-stage bimanual interactions. Existing real-robot benchmarks mainly focus on short-horizon rigid-object tasks and offer limited coverage of long-horizon deformable manipulation. We introduce FolDeX, a physical-world benchmark built entirely from real-robot data, with garment folding as its primary task. Since real-robot data collection is costly, FolDeX studies how heterogeneous physical experience can be reused efficiently. The benchmark is organized around four research axes: leveraging human intervention and recovery data collected during deployment; transferring data across tasks, including across garment categories and from rigid to deformable-object manipulation; reusing data across scenes with changes in lighting, background, and layout; and transferring data across robotic embodiments. FolDeX provides 2,000+ hours of real-robot data spanning 20+ tasks and 10+ embodiments. We also establish a fair real-robot evaluation platform for externally submitted policies, with standardized tasks, held-out physical objects, controlled initializations, and a unified execution protocol. The platform is publicly accessible at https://ai.midea.com/#/fold-challenge. We hope FolDeX will serve as a unified testbed for heterogeneous real-robot data reuse and reliable long-horizon deformable manipulation.

Wed 9 SeptRobotics
The gist
Robots often struggle to handle soft, flexible objects in the real world, especially when tasks require many steps. The authors introduce FolDeX, a large collection of real robot data focused on folding clothes, which is a challenging long task involving soft items. This benchmark helps test and improve robot skills by allowing researchers to reuse data from different tasks, objects, and setups. It includes over 2,000 hours of robotic activity and a public evaluation platform for testing robot policies on consistent tasks and conditions.
Open 2609.10243v1

Robot adapts shared control to human behavior for better cooperation

Adaptive Shared Control with Online Bounded-Rational Human Behavior Estimation

Abstract: This work considers adaptive shared human-robot control for nonlinear control-affine systems, where the assumption of a fully rational human is relaxed and the robot adapts its assistance to observed boundedly rational human behavior. We use a level-k bounded-rationality model of the two-player game to construct a finite bank of candidate human and robot policies through alternating best-response computations, with the associated value functions and policies approximated using adaptive dynamic programming. During the shared-control interaction, state-transition residuals compare the measured system evolution with the trajectories predicted by the candidate human policies. The residuals are accumulated using a forgetting factor and mapped to a probabilistic human-behavior model over the finite candidate bank. Rather than selecting a single candidate or averaging stored robot policies, the robot computes a distribution-aware one-step best response by minimizing an expected cooperative cost over the complete estimated human behavior distribution. For a quadratic terminal-value approximation and Euler state propagation, this response admits a closed-form solution expressed in terms of the expected human input. The proposed methods are evaluated in simulations of a benchmark nonlinear system stabilization task, and of a planar manipulator shared control setup. The reported results show decreasing Kullback-Leibler divergence between the estimated and simulated human behavior distributions, and a lower accumulated running cost for the robot agent over the shared control interaction period, than the maximum-probability and probability-weighted alternative policies baseline.

Wed 9 SeptRobotics
The gist
Robots and humans sometimes work together to control machines, but humans don’t always act perfectly logically. The authors developed a way for a robot to guess what kind of thinking a human partner might have during control, using patterns of past actions. Instead of trusting just one guess, the robot considers a range of possible human behaviors to decide how to assist. This approach helps the robot provide better support, shown in tests where the robot’s actions matched the human’s style more closely and reduced errors.
Open 2609.10215v1

Robotic hand learns to assemble two parts using one hand

Assembling Two Parts in One Hand

Abstract: A hallmark of human dexterity is the cooperative use of fingers, where different fingers take on distinct yet coordinated roles to accomplish fine manipu- lation, such as capping a pen with the hand that holds it. We study this finger-level coordination through in-hand assembly: mating two rigid objects within a single dexterous hand, with no second arm and no fixture. We present a reinforcement learning formulation to solve this problem in a unified framework, which is driven by a goal relative pose between the two parts. Finger coordination is shaped by a function-based auxiliary reward and regularized toward a single human reference pose, while domain randomization and a fusion of historical proprioception and object observation confer robustness to occlusion-induced estimation noise. The same recipe solves three different assembly tasks (Bottle, Syringe, and Marker). Trained purely in simulation, the policies transfer zero-shot to hardware with a single camera, demonstrating robustness to state-estimation errors caused by oc- clusion. Our experiments also reveal that in-hand assembly places demands on hand morphology and can serve as a benchmark for modern robotic hand systems. Videos and code are available at https://ltbgbird.github.io/in-hand-assembly-page/.

Wed 9 SeptRobotics
The gist
Putting two objects together using only one hand is something humans do with ease, but it’s hard for robots. The authors taught a robot hand to hold and fit two objects together without needing help from another arm or tool. They used a kind of trial-and-error learning method, trained the robot in a computer simulation, and then successfully tested it on a real robot hand. The method works even when the robot can’t perfectly see the objects it’s holding, showing it can handle tricky situations like occlusion.
Open 2609.10137v1

Automatic camera calibration improves image selection and parameter choice

Automatic Reproducible Camera Intrinsic Calibration

Abstract: Accurate camera intrinsic calibration is fundamental to robot perception, and the accuracy depends on the quality of the collected images. However, existing target-based calibration methods often require the practitioner to manually filter out high-quality images and to specify an appropriate radial distortion order. This paper presents a fully automatic intrinsic calibration pipeline that determines both from the collected data. We adopt an iterative rejection scheme that estimates parameters on a candidate image set and removes views whose mean residual exceeds a multiple of the median. Crucially, this process runs independently under each candidate distortion order, so that the retained image set is consistent with the residual scale of that order. Further, the distortion order is selected on held-out images, with the intrinsics and distortion fixed and only the board pose re-estimated, ensuring that an added coefficient is supported by independent observations. Finally, we integrate both steps into an interactive calibration tool that supports full-pipeline data inspection and parameter estimation. Experiments on our own camera data and five public real-world datasets show that image filtering reduces the held-out reprojection error by 25\%, the order selection further by 5\%, achieving the lowest held-out mean among four compared configurations without manual image selection. We will release the code and data to facilitate future research.

Wed 9 SeptRoboticsComputer Vision and Pattern Recognition
The gist
Camera calibration helps robots understand what they see by figuring out how the camera lens changes the image. Usually, people must pick good photos and guess how complex the lens distortion is. This paper presents a method that does both steps automatically by testing and rejecting bad photos and choosing the right complexity for distortion based on new data. The authors also made an interactive tool and showed that their method lowers errors significantly compared to not filtering images or choosing distortion by hand.
Open 2609.10082v1

Video plans improve versatile robotic hand manipulation in simulation

Grounding Generated Video Plans in Simulation Towards Versatile Dexterous Controllers

Abstract: Generated hand-object interaction (HOI) videos provide a controllable way to propose manipulation motions. Simulation-based HOI tracking can translate such kinematic references into feasible low-level control, but its scalability is limited by the lack of reliable reference motions. We therefore combine generated videos with simulation-based HOI grounding: during training, generated videos provide diverse motion references for learning a multi-object, multi-trajectory HOI tracker, and at deployment, the video model produces motion plans that are executed by the learned tracker. In particular, we propose a method that enables scalable reference generation by HOI reconstruction with minimal manual intervention and successfully grounds more than 1,500 generated videos in simulation, achieving success rates over 25 percentage points higher than those of baselines during simulation-based training. In real-world closed-loop experiments, it achieves diverse grasps, including functional grasps, non-prehensile manipulation, and post-grasp object-pose tracking. Videos and code are available at https://boyuan-an.github.io/GALATEA/.

Wed 9 SeptRobotics
The gist
Controlling robot hands to manipulate objects is hard because they need precise movements. The authors show how computer-generated videos of hand-object interactions can be used to teach robots by translating these videos into control plans. Their method makes it easier to create many examples and helps robots perform various grasp and object-moving tasks better than previous methods. It works both in simulation and on real robots, enabling more versatile and reliable robotic manipulation.
Open 2609.10050v1

Belief state engine improves planning in limited view environments

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

Abstract: Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation can collapse their uncertainty onto the wrong hypothesis. Policies drift as the history grows. We trace these symptoms to a common structural cause. An LLM agent, as commonly deployed, is a history-conditioned policy with no explicit belief over hidden state. We propose an architectural fix. The Belief-State Engine (BSE) is an inference module placed outside the LLM. It maintains a Bayesian posterior over the latent states of a given POMDP (Partially Observable Markov Decision Process) model, and at each decision step it exposes only that posterior to the LLM. The raw action-observation log is not shown. We set out a minimal four-axiom specification of what a belief-consistent internal state must satisfy, and prove that the LLM paired with the BSE is a sound Markov policy on the belief MDP induced by the underlying POMDP. It therefore inherits the Bellman optimality guarantees of classical POMDP theory, provided the LLM is never exposed to the raw history. We evaluate the architecture on the Tiger POMDP and a red-team attack-graph task, against six baselines: a reactive LLM, Chain-of-Thought, ReAct, a natural-language belief tracker, QMDP, and POMCP. Across both domains, the BSE-augmented agent improves task return, belief calibration, and decision consistency. Ten targeted ablations isolate the contribution of each architectural choice confirms that the effect is not specific to any one model. Code, environment specifications, prompt templates, and seed logs accompany this paper.

Wed 9 SeptArtificial IntelligenceMachine LearningRobotics
The gist
Large language models (LLMs) can follow instructions to perform tasks, but they struggle when they can't see the whole situation clearly. The authors found that this happens because these models don’t keep a clear mental map of what might be hidden. They built a tool called the Belief-State Engine (BSE) that keeps track of what is likely true based on past actions and observations. By only showing the LLM this clear summary instead of all past details, the system plans better and makes smarter decisions in tricky situations.
Open 2609.10036v1

Using symmetry boosts learned 3D motion planning success rates

What Symmetry Buys a Learned Motion Planner

Abstract: Learning-based motion planners pay at training what classical planners pay per query. Trained in world coordinates, they relearn the same motion at every position and orientation. Existing work restores the missing rigid-body equivariance in the training data, in the inference operator, or in the weights, and each carries a cost. We ask how much of that equivariance the planning query supplies for free. A start s and a goal g determine a frame in closed form, with origin at their midpoint and first axis along g-s. Expressing trajectory and obstacles in that frame removes three translations and two rotations of SE(3), at initialisation, for one cross product per query and with no constraint on the architecture. A single rotation about the start-goal axis remains, and no continuous rule removes it. On a cluttered 3D benchmark, holding architecture, data and budget fixed, the frame raises the held-out collision-free rate from 14.60% to 51.10%, where a straight segment from start to goal scores 15.6% and the world-frame model does not beat it. We build all three mechanisms for the residual rotation and each is worth under a point, though the equivariant backbone reaches any given level two to three times sooner. What the representation supplies therefore dominates what any mechanism enforces, and the standard diagnostic does not see the difference: two models with indistinguishable non-equivariance residuals differ by 28 points. Calibrated against a non-symmetry intervention, the frame is not even the largest effect available, since local geometry is worth +40.0 where the frame is worth +36.5.

Wed 9 SeptRobotics
The gist
Planning paths for moving objects in 3D space is hard because the same motions must be relearned many times. The authors show that by smartly aligning the problem based on the start and goal positions, a motion planner can automatically ignore many unnecessary orientations and positions. This leads to much better success in finding collision-free paths compared to traditional methods. They also find that adding more complex symmetry fixes adds little beyond this alignment step.
Open 2609.10033v1

Axon improves ROS 2 robot communication with shared memory and quantum keys

AXON: A ROS 2 RMW with Shared-Memory/QUIC Transport and QKD/ML-KEM Key Establishment

Abstract: Robot Operating System 2 (ROS 2) standardizes application code against a middleware interface (RMW) whose reference implementations are built on the Data Distribution Service (DDS). We present AXON, an alternative ROS 2 RMW implementation that separates transport policy by deployment scope. A Rust core and C++ adapter use POSIX shared-memory rings for same-host communication, QUIC for remote communication, and a daemon for discovery and graph synchronization. We then describe two fail-closed TLS 1.3 key-establishment configurations for remote traffic. The classic configuration offers only the hybrid X25519MLKEM768 group, preventing negotiation of a classical-only group. The qkd configuration imports a 256-bit key obtained through the ETSI GS QKD 014 API as a pairwise external PSK and offers no Diffie-Hellman group. Its default messages10 strategy additionally protects remote application messages with AES-256-GCM, rotating KME material after ten outgoing messages and using a fresh nonce per envelope; session relies on QUIC protection alone. The external-PSK path requires a narrow extension to rustls, now bundled with AXON. We define the threat model, distinguish peer authentication in the two configurations, and delimit the implementation-level validation from ROS 2 conformance, comparative performance, and physical-QKD validation.

Wed 9 SeptRobotics
The gist
ROS 2 is software used to help robots talk to each other and to computers. The authors created AXON, a new way for ROS 2 to send messages more efficiently by using shared memory when on the same machine and a fast internet protocol for remote communication. They also added a special security method that uses quantum key distribution to protect messages from being hacked. This setup helps keep robot communications fast and secure, especially over networks.
Open 2609.10024v1

RoboDrop improves robot learning by filtering training data errors

RoboDrop: Curating VLA Post-Training Data via Local Gradient Compatibility

Abstract: Vision--language--action (VLA) models acquire broad generalization through large-scale pretraining, yet adapting them to a new task and robot embodiment still requires post-training on newly collected data. Unlike pretraining, post-training targets task- and embodiment-specific adaptation, making it particularly sensitive to data quality. In practice, collected robot datasets often contain heterogeneous errors, including execution mistakes, sensor drift, and timestamp misalignment, which can impair post-training and policy performance. Manual inspection is costly, while existing data-cleaning methods are typically tailored to particular corruption types. To address these challenges, we introduce \textsc{RoboDrop}, a data-curation framework that audits supervision using local gradient compatibility measured along the training trajectory as a proxy for its effect on post-training performance. During a one-epoch warm-up run, RoboDrop scores each candidate sample online by comparing its gradient with those of task-semantic and visually matched validation samples. The resulting sample scores are aggregated at the episode level, and a simple automatic post-processing rule converts them into filtering decisions. We evaluate RoboDrop on controlled observation--action corruptions, naturally suboptimal demonstrations in simulation, and real-robot datasets containing non-expert collection errors. Across these settings, RoboDrop more accurately distinguishes unreliable demonstrations than prior methods, while post-training on the curated data consistently yields stronger downstream policies, with average real-robot rollout success rising from $35.0\%$ to $67.5\%$. These results establish training-trajectory-aware, context-conditioned supervision auditing as an effective approach to robust VLA post-training.

Wed 9 SeptRobotics
The gist
Robots learn new tasks by training on data collected from their own actions, but this data often has mistakes like sensor errors or timing problems that confuse the learning process. The authors present RoboDrop, a tool that automatically checks training data for quality by watching how each example influences learning progress. It scores and filters out bad data, helping the robot learn more effectively. Testing shows RoboDrop improves the robot’s success rate by removing problematic examples from the training set.
Open 2609.10021v1

Hallucination-aware world model improves general robot manipulation success

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

Abstract: Generalist robot policies have demonstrated strong generalization across robotic manipulation tasks, yet their success rates remain limited in com- plex long-horizon scenarios. Recent methods improve Visual-Language-Action (VLA) policies through online reinforcement learning on real robots, but such training relies on costly physical interactions, suffers from low sample efficiency, and may introduce hardware and safety risks. World models offer a promising alternative by enabling policy optimization with imagined rollouts. However, long-horizon rollouts generated by world models often suffer from prediction hal- lucinations, producing biased state transitions that can mislead policy learning. To address this issue, we propose Hallucination-aware World Model-based Pol- icy Optimization (HaWMPO), a closed-loop reinforcement learning pipeline for VLA policy post-training with world models. Specifically, HaWMPO introduces an action-conditioned hallucination-aware model to estimate the reliability of gen- erated image sequences, and incorporates hallucination scores into group relative policy optimization through a Reward-Soft mechanism, suppressing unreliable ac- tion chunks during training. On the LIBERO benchmark, HaWMPO achieves the best average success rate, with gains of 15.0% over the base model and 2.8% over the strongest baseline; real-world experiments on a G1 robot further validate its effectiveness, raising the average success rate on two manipulation tasks from 67.5% to 80.0%.

Wed 9 SeptRobotics
The gist
Robots that can do many tasks often struggle with long and complex sequences of actions. The authors found that training robots using imagined experiences called world models can help, but these imaginations sometimes have mistakes, called hallucinations, that confuse the robot. They created a new method called HaWMPO that watches out for these hallucinations and lowers their influence on learning. This approach improved robot success rates significantly in tests and real-world tasks.
Open 2609.09941v1

Vision language models improve robot actions with better time frequency coding

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

Abstract: Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is essentially a short multivariate trajectory, but inside these models it is a sequence of generic per-timestep hidden tokens decoded by a linear head. This under-serves two motion structures. First, frequency: a chunk superimposes a smooth global trend and fine corrective motion across time scales, and a single token entangles them. Second, cross-phase geometry: motions of different phases (reach, contact, grasp adjustment, settling) unfold along very different, near-orthogonal directions in representation space, yet are tightly related for the task and arise across the time axis. Dot-product attention scores alignment by an inner product, so it favors aligned tokens and is least sensitive near orthogonality, leaving such relationships for the network to recover through a detour. We introduce Time-Frequency Geometric Cross-Attention (TFGCA), a drop-in module repairing both blind spots. TFGCA uses a per-dimension learnable stationary wavelet transform to decompose the action chunk into time-frequency tokens, and each time token retrieves information from them via a cross-attention that fuses the dot product (similarity) with the wedge-product magnitude (sensitive to near-orthogonality) through a learnable weight. A zero-initialized residual reproduces the base behavior at initialization, so it can be dropped onto a pretrained VLA and fine-tuned jointly. Relative to the same-source base, TFGCA improves in-distribution LIBERO by +1.5 on average, the OOD LIBERO-Plus by +6.3, the randomized average under RoboTwin domain randomization by +28.5, and the overall success rate on three real-robot AgiBot A2 tasks by +11.67 points, with larger gains out of distribution.

Wed 9 SeptArtificial IntelligenceRobotics
The gist
Robots need to plan many steps of movement at once, but current models treat these as simple sequences of tokens, missing details in how actions change over time. The authors designed a new method that breaks down movements by their smooth trends and quick corrections and also better connects different action phases that usually look very different to the model. This method combines standard similarity with a new way to detect very different but related motions, improving robot task performance especially when facing new situations. Their approach can be added to existing robot models without starting from scratch.
Open 2609.09925v1

Humanoid robots adapt visual feedback for whole-body tasks

ViBe: Visual Behavior Adaptation for Perceptive Humanoid Whole-Body Control

Abstract: Motion tracking provides a scalable recipe for humanoid whole-body control. By design, the resulting trackers lack exteroceptive feedback hence reacting to the environment remains the responsibility of a higher-level planner. Existing perceptive controllers train geometry-only encoders from scratch, trading semantics for sim-to-real ease, and typically rely on teacher-student distillation for a task of interest. We present ViBe, a post-training framework for adapting motion trackers to perceptive control tasks. We leverage pre-trained visual encoders with a multi-query extractor module to learn task-relevant perceptive feedback. This feedback is grafted onto the tracker's input via low-rank adapters, enabling parameter-efficient fine-tuning. Given a task reward and a reference dataset, this modular controller can be adapted directly via policy optimization. Across four tasks, ViBe shows zero-shot sim-to-real transfer spanning perceptive walking on curbs and parkour, Repose Cube, omni-object loco-manipulation, and dodgeball, with visually robust performance across outdoor, low-light, and RGB distractor conditions. Finally, we solve a goal-oriented Repose Cube task with a deliberately simple planner, demonstrating the efficacy of perceptive controllers, adapted by our approach.

Wed 9 SeptRobotics
The gist
Humanoid robots usually follow pre-recorded motions but struggle to react to changing environments because they lack visual feedback in their control systems. The authors present ViBe, a way to add vision-based sensing to existing motion trackers by using pre-trained visual encoders and fine-tuning them efficiently for specific tasks. This enables robots to better handle activities like walking over curbs, parkour, object manipulation, and even playing dodgeball, while transferring these skills from simulation to the real world. The approach was tested in various conditions including outdoors and low light, showing robust performance without needing a complex planner.
Open 2609.09918v1

Real-time adaptation links real deformable object data to simulations

RealSimLoop: Online Real-to-Sim Adaptation via Differentiable Reduced-Order Simulation with Vision Feedback

Abstract: Real-world observations of deformable objects are often sparse or surface-level, while downstream tasks require hidden physical quantities such as internal deformation, stress fields, and interaction forces. Physics-based simulation can recover these quantities, but online real-to-sim adaptation remains challenging due to costly full-space optimization, limited feedback, and time-varying material properties. To address these challenges, we propose RealSimLoop, a differentiable framework for online real-to-sim adaptation using vision data as physical feedback. Our approach achieves quasi-real-time performance by executing differentiable simulation within a reduced-order neural subspace, drastically accelerating the optimization loop. We couple this efficient dynamics model with differentiable rendering, enabling direct gradient backpropagation that leverages high-fidelity pixel data to refine physical parameters such as material stiffness. Furthermore, by employing a sliding-window objective function, RealSimLoop enables robust online adaptation, allowing the system to track time-varying material properties and effectively bridge the real-to-sim gap arising from model reduction or unmodeled dynamics. Extensive experiments demonstrate that our method outperforms conventional offline methods, and we validate the framework's versatility in downstream applications, including external force prediction and 3D stress field reconstruction with novel view synthesis.

Wed 9 SeptGraphicsComputer Vision and Pattern RecognitionRobotics
The gist
Simulations can show hidden details about deforming objects, like their stress or forces, but usually don’t match real-world changes quickly. The authors created RealSimLoop, a method that uses camera images to quickly adjust simulation settings to match real objects, even as those objects change over time. It does this efficiently by simplifying the simulation and using image data directly to update the model. Their tests show better results than traditional methods, with uses like predicting forces and visualizing internal stresses from new viewpoints.
Open 2609.09828v1

InstantMimic speeds up physics learning for realistic character control

InstantMimic: A High Performance System for Learning Physics-based Skills in Seconds

Abstract: Physics-based character control is a long-standing challenge in computer graphics and robotics, requiring policies that satisfy complex dynamics while producing realistic motion. Recent Deep RL approaches, particularly imitation learning methods such as DeepMimic, have had broad impact beyond animation, influencing robotics by enabling agile and expressive behaviors. While these approaches achieve impressive results, they remain computationally inefficient to train in practice. Despite GPU-accelerated simulation, we find that end-to-end pipelines often underutilize hardware due to overheads outside the physics solver, caused by fragmented GPU kernels and CPU memory access in the critical path. We present InstantMimic, a system that addresses these inefficiencies by making the entire training loop GPU-native. Built on a GPU-native physics backend, our unified pipeline integrates simulation, environment computation, policy inference, and policy updates within a single execution flow. As a result, InstantMimic reduces training time for diverse physics-based skills to a few seconds and makes LLM-agent-driven hyperparameter search practical.

Wed 9 SeptGraphicsRobotics
The gist
Making computer characters move realistically using physical rules takes a lot of computing time, slowing down how quickly these skills can be learned. The authors found that existing methods waste computing power, especially when using both CPUs and GPUs inefficiently. They created InstantMimic, a tool that runs the whole learning process entirely on the GPU, cutting training time to just a few seconds. This makes it easier and faster to teach characters complicated physical skills and allows for quick experimentation with settings.
Open 2609.09821v1

PccDiffuser plans multiple safe paths for soft robots in cluttered spaces

PccDiffuser: Multi-solution Motion Planning for Continuum Robots

Abstract: We present the PccDiffuser, a conditional diffusion framework for continuum robots that learns a multimodal distribution over complete configuration-space paths and samples multiple candidate solutions in parallel, which are subsequently converted into an executable trajectory by time allocation considering actuator constraints. Under the piecewise constant-curvature model, we use exponential co-ordinates to describe the robot kinematics, and use graph neural network to encode a variable number of environment obstacles. Analytical differential kinematics is incorporated in the denoising process to improve terminal accuracy and whole-body clearance. On a mixed test set comprising workspace with zero to four obstacles, PccDiffuser achieved a success rate of 91\%. Compared with existing sampling- and optimisation-based benchmarks, it delivered both a higher success rate and greater computational efficiency, with the latter advantage becoming more substantial when sampling more candidate solutions. Experiments on a three-section tendon-driven continuum robot further demonstrate consecutive planning, multi-solution planning, and whole-body obstacle avoidance.

Wed 9 SeptRobotics
The gist
Planning how soft, bendy robots move through spaces with obstacles is hard because these robots can curve in many ways. The authors introduce PccDiffuser, a new method that learns many possible safe paths simultaneously and picks the best ones quickly. It uses math models for how the robot bends and a type of neural network to understand obstacles. Tests show it finds good paths more often and faster than older methods, even with multiple obstacles.
Open 2609.09745v1

Actionsplice enables instant action updates in video world models

ActionSplice: In-Flight Action Editing for Interactive World Models

Abstract: Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition future solver evaluations on a state produced under the previous action, or trigger rollback that repeats completed evaluations. We introduce ActionSplice, an inference framework that formulates this problem as Counterfactual State Transport (CST). A lightweight corrector transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. The retargeting variant $\mathrm{CST}*{R}$ updates the entire active chunk, while the temporal-splicing variant $\mathrm{CST}*{T}$ preserves a temporal prefix and updates only the suffix. Across minWM-Wan Action2V and HY-WM1.5, $\mathrm{CST}*{R}$ reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. $\mathrm{CST}*{T}$ reduces suffix LPIPS by 56.1% and 77.5%, respectively, while providing $2.73\times$ and $1.69\times$ pixel-ready speedups over waiting. Under the HY-WorldPlay protocol, $\mathrm{CST}_{R}$ obtains a PSNR of 25.66 dB, an SSIM of 0.6902, and an LPIPS of 0.1337 against the original rollout.

Tue 8 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
When making video models that predict future scenes based on actions, updating an action partway through is tricky and often slow. The authors introduce ActionSplice, a method that lets these models quickly adjust to new actions without redoing all the previous work. It uses a lightweight corrector to shift the model’s state to match the new action and allows smooth continuation. This approach improves accuracy and speeds up video prediction when actions change during generation.
Open 2609.08230v1

Robot learns efficiently from targeted demonstration requests

DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning

Abstract: A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly that. Interactive imitation learning takes a step in that direction by letting a policy practice on its own and calling an expert when it goes wrong. Existing methods decide when to interrupt the learner. A further 2 decisions are left to whichever episode happened to trigger the interruption: which failure to correct, and where the demonstration should start. This paper is a first attempt at making both of them deliberately. DISEIL (Demonstration dIstillation for Sample-Efficient Imitation Learning) marks each failed episode at the step where the policy first becomes unreliable, represents that moment with a geometric descriptor, and groups the failures into recurring failure modes. A vision-language model and a language model read the selected mode and write a request for the next demonstration, and a store of task constraints checks that the request can be carried out before any expert time is spent. No model produces a robot action. Across 5 simulated tasks under state and image observations, changing only what the expert is asked for gives the highest mean held-out success rate in all 10 settings, with a tie in 1, and the margin is widest at the smallest budget we tested. The scope is narrow: a single round of practice at a time, in simulation, with experts that are mostly scripted. The longer-term aim is a learner that also tracks what its demonstration set already covers, and that asks a human teacher for the missing behavior in proportion to the effort each request costs them.

Tue 8 SeptRoboticsArtificial IntelligenceMachine Learning
The gist
Robots can learn new tasks faster if they know exactly what parts they struggle with and ask for help on those parts. The authors developed a system that watches where a robot first makes mistakes during practice and groups similar errors. Then, using language and vision models, it generates specific requests for demonstrations to fix those errors, avoiding wasted expert time. This method improved learning success in various simulated tasks, especially when only a few demonstrations were allowed. The work focuses on a single practice round in simulation with mostly scripted experts.
Open 2609.08123v1

InfluenceField improves prediction of visual intervention effects in language AI

InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling

Abstract: Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers. We introduce InfluenceField, an intervention-aware latent field inserted between the visual encoder and language decoder. It lifts patch features into a continuous spatial representation, propagates directed influence over multiple steps, and predicts local intervention effects through a shared transition operator. Training jointly optimizes language modeling, cross-environment invariance, counterfactual rollout supervision, and structural regularization. For a nonlinear finite-basis population model, we show that target-aligned interventional supervision, together with a one-step separation condition on the transition, restricts admissible representations to within-location reparameterizations, so that the directed dependency graph of the full transition is recovered exactly. A linear specialization gives an exact partial-coverage characterization and a finite-loss stability bound, and the field analysis derives the spatial profile of coefficient interventions together with a shared-channel calibration result. On CausalVQA, InfluenceField improves overall accuracy over its backbone by 13.1 percentage points, with the largest gains on the planning and hypothetical categories. Capacity-matched baselines and structural controls attribute the gains in robustness and factual-counterfactual consistency to the causal objectives rather than to added capacity.

Mon 7 SeptMachine Learning
The gist
Multimodal AI models that work with both images and language often find it hard to understand how making changes to specific parts of an image affects answers about the scene. The authors propose InfluenceField, a new component that sits between the image understanding part and the language generation part, capturing how changes at one point in an image influence others. This method learns to predict the effects of local changes by modeling directed influence across space and time, improving accuracy especially on complex tasks like planning or imagining what-if scenarios. Their approach also provides a way to identify causal relationships within the model's internal representations.
Open 2609.07874v1

Vision language action models improve with aligned few shot demos

ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models

Abstract: Vision-Language-Action (VLA) policies are commonly adapted to new manipulation settings through additional gradient updates, which limits rapid deployment when task-specific data or compute is scarce. We present ICI-VLA, a training and retrieval framework that equips a text-action VLM with few-shot test-time adaptation through in-context demonstrations. Unlike mainstream VLA designs based on action-specific multimodal fusion, ICI-VLA retains the native text-generation interface. ICI-VLA updates its parameters only during offline training; at inference, the policy remains fixed and conditions action generation on retrieved micro-demonstrations. The framework decomposes long trajectories into short, semantically labeled examples and trains an RD-Encoder with positives mined by Dynamic Time Warping (DTW), aligning the retrieved context with the phase and geometry of the current subtask. We further introduce Target Action Masking, a context-corruption objective designed to reduce direct action copying and increase reliance on the current observation. ICI-VLA reaches average success rates of 97.7% on LIBERO and 60.4% on RoboTwin 2.0, exceeding the highest reported baseline average on RoboTwin 2.0 by 19.3 percentage points. It also achieves 83.2% across four physical tasks. These results indicate that a fixed VLA policy can benefit from conditioning on spatiotemporally aligned demonstrations at test time.

Mon 7 SeptRobotics
The gist
It can be hard for robots to quickly learn new tasks without lots of practice or computing power. The authors built a system that helps a robot follow instructions better by showing it short video examples closely matched in time and space to what it’s currently doing, without changing the robot’s learned skills on the fly. Their method breaks long tasks into smaller parts and uses a smart way to find matching demonstration clips that guide the robot’s actions. This approach works very well on several robot testing environments, beating previous methods and even succeeding on real robot tasks.
Open 2609.07581v1

Flying humanoid robots walk on ceilings with smoother thrust control

Anti-Gravity Walking by a Flying Humanoid Robot via Thrust-Rate Input Whole-Body Model Predictive Control

Abstract: Flying humanoids are expected to perform tasks in diverse environments, while their existing locomotion is mainly limited to aerial flight and ground walking. The capability to move in complex three-dimensional space can greatly expand their application range. For such walking motion on ceilings and similar anti-gravity environments, whole-body MPC is effective. However, the discontinuous changes in dynamic structure accompanying contact switching during walking can induce thrust spikes, resulting in control instability. Therefore, in this work, we propose and implement a real-time whole-body MPC framework for anti-gravity bipedal walking. First, we formulate whole-body MPC using the time derivative of thrust, namely thrust-rate, as the control input. This formulation guarantees continuity of the thrust trajectory during contact switching while preserving the sparse structure of the optimal control problem for fast computation. Second, we address the lack of natural support forces in anti-gravity environments. We introduce lower bounds on the foot-normal component of the contact force, and smoothly transfer them during the doublesupport phase. Finally, we implement the proposed framework and demonstrate anti-gravity walking by a flying humanoid through simulation and a hardware experiment. To the best of our knowledge, this is the first demonstration of multi-contact whole-body MPC for a transformable aerial robot and walking by a flying humanoid beyond the ground.

Mon 7 SeptRobotics
The gist
Flying robots that look like humans usually either fly or walk on the ground. This work helps these robots walk on ceilings or upside-down surfaces by improving how their controls handle forces. The researchers designed a control system that changes the force output smoothly even when the robot's feet switch contact points, avoiding sudden spikes that could cause problems. They also added special constraints to mimic support forces when the robot walks upside down. They tested these ideas with simulations and an actual robot, showing it can walk in anti-gravity conditions.
Open 2609.07544v1

Proprioception improves robot assembly from simulation to reality

Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining

Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained contact. Although simulation-based reinforcement learning offers a scalable training paradigm, discrepancies in visual observations, contact dynamics, and force/torque (F/T) measurements often limit policy transfer. We observe that proprioception is comparatively consistent across domains because calibrated joint positions and consistently computed joint velocities align closely between simulation and hardware. Based on this observation, we present PACE (Proprioception-Anchored Cross-Modal Encoder), which supervises temporal visual and F/T representations by predicting proprioceptive state transitions. Static domain-specific factors, including lighting, texture, and sensor bias, contain little information about joint motion; the proposed objective therefore encourages the encoder to suppress these factors while retaining task-relevant motion cues. Policies trained on frozen PACE features are deployed on hardware without real-world fine-tuning or object-pose tracking. Across four contact-rich assembly tasks, PACE attains an average real-world success rate of 93.3\% and only a 2.7-percentage-point sim-to-real drop, meanwhile remaining robust to perturbations that substantially degrade pose-based and learned-fusion baselines.

Mon 7 SeptRoboticsArtificial Intelligence
The gist
Contact-rich assembly tasks, like snapping parts together, are hard for robots because they need super precise movements and must understand forces while touching objects. The authors found that a robot's sense of its own joint positions and movements, called proprioception, stays consistent between simulations and the real world. They created a system named PACE that uses this reliable proprioceptive data to better learn how visual and force information relates to movement. Robots trained with PACE worked well in real assembly tasks without extra tuning, achieving over 90% success and staying strong even when things changed unexpectedly.
Open 2609.07534v1

Robot language skills improve with Greek added to bilingual training

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

Abstract: Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rather than translation. Several plausible instruments produce false conclusions: a color-histogram metric rewards noise, a single-goal benchmark scores 84.6% under correct Greek and 82.6% under deliberately wrong instructions, training loss fails to predict Greek success, and single-run comparisons are dominated by seed variation. On a discriminative ninety-task suite with three seeds per arm, a multilingual text tower without Greek demonstrations remains at its wrong-instruction floor, while Greek-only training exceeds its control by at most 2.7 points. Bilingual training yields a consistent 6.7-7.1 point margin over its control and reaches about two fifths of English performance. The policy also overfits the translator's phrasing; training on seven phrasings per task approximately halves this penalty. Warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance. The results support two practical requirements for low-resource robot-policy localization: build a guaranteed null before trusting a metric, and replicate low-resource-language results across seeds.

Mon 7 SeptRoboticsArtificial Intelligence
The gist
Robots usually learn commands only in English because there aren't many training examples in other languages. The authors tried teaching a robot to understand Greek instructions using machine-translated commands without changing its design. They found it’s hard to measure if the robot really understands Greek, and some usual tests give wrong answers. Training with both English and Greek together helped the robot understand Greek better than just Greek alone, reaching about 40% of its English ability. They also found the robot learns specific ways of saying things too closely, but teaching it many ways to say the same command helps fix that.
Open 2609.07470v1

Legged robots separate walking noise from environment sounds without labels

Open-Set Ego-Noise Separation for Legged-Robot Audition via Annotation-Free Adaptation and Pretrained-Model Transfer

Abstract: This paper proposes an open-set ego-noise separation framework for legged-robot audition via annotation-free adaptation and pretrained-model transfer. The framework removes robot-specific ego-noise while preserving environmental sounds whose classes are not specified in advance. Acoustic sensing provides cues about a robot's surroundings beyond the visual field, but walking-induced ego-noise from footstep impacts, joint-backlash rattling, and motor noise severely contaminates the recordings. The framework first uses RecurGraph to select ego-noise-dominant clips from the unlabeled recordings by aggregating clip embeddings into an embedding centroid and propagating scores over an audio-embedding graph. The selected clips are mixed with diverse environmental sounds from a large-scale sound-event dataset to provide paired mixture--target supervision for open-set separation. Transfer-DiT then adapts a general-purpose zero-shot neural separator to achieve high-fidelity open-set ego-noise separation for the target robot. Experiments with bipedal and quadrupedal robots show reliable clip selection and improvements in separation quality and downstream task performance over baseline separators. These results demonstrate the feasibility of annotation-free adaptation without separately recorded ego-noise-only data or manual clip-level annotations.

Mon 7 SeptRobotics
The gist
Legged robots have trouble hearing what’s around them because their footsteps and motors make lots of noise. The researchers created a way for robots to learn and remove their own walking noises from sound recordings without needing people to label the noise beforehand. They use a method that finds parts of recordings mostly made up of robot noise and then teach a sound separation model using those plus extra environmental sounds. This approach works well on robots with different numbers of legs and improves their ability to understand sounds in noisy situations.
Open 2609.07440v1

Robot arm position affects how close people let it get

CALM: Configuration-Aware Human Intervention Boundaries During Robot Approach

Abstract: How robot body configuration shapes human intervention during approach remains underexplored. We conducted a within-participants study with 41 participants, measuring final stopping distance, subjective comfort, and exploratory eye-tracking responses across four humanoid arm configurations and two spatial scales. Full forward arm extension increased stopping distance by approximately 31-36 cm relative to arms-down. Spatial scale primarily affected comfort and pupil responses without a detectable stopping-distance shift. We introduce the Configuration-Aware Limit Model (CALM), which translates stopping-distance distributions into configuration-dependent population-coverage boundaries. Estimated boundaries at 80% coverage ranged from 0.88 to 1.47 m. In an illustrative one-dimensional planning analysis, reconfiguration enabled a 1.10 m approach goal that was unreachable with arms remaining fully extended under the same nominal pointwise 20% intervention-probability constraint. These findings support treating body configuration as a planning variable while distinguishing physical safety, behavioral intervention, and subjective cost.

Mon 7 SeptRoboticsHuman-Computer Interaction
The gist
People feel differently about how close a robot can come based on how it holds its arms. The study found that when a humanoid robot's arms were fully extended forward, people stopped it from coming as close compared to when its arms were down by its sides. The researchers created a model called CALM that predicts how close people will allow a robot to get based on its arm position. This helps robots plan their movements to avoid making people uncomfortable or feeling unsafe.
Open 2609.07430v1

OpenWAM delivers versatile pretrained world and action models for robotics

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

Abstract: World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

Mon 7 SeptRobotics
The gist
Robots and AI can learn how to understand and act in the world by combining knowledge from videos with hands-on experience. The authors created OpenWAM, a modular platform that lets researchers test different ways to build and train these models to see what works best. They discovered key design principles that improve how well these models learn and generalize. Using these insights, they built OpenWAM-α, pretrained on thousands of hours of human and robot video data, which performs well in both simulated environments and real robots. The whole system and tools are openly shared to support future robotics development.
Open 2609.07398v1