Papers for

robot software engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Sample simulate and select improves text to robot motion without training

Sample, Simulate, Select: Physics-in-the-Loop Text-to-Motion for Humanoids Without Training

Abstract: Text-to-motion models generate plausible human motion but do not model a robot's dynamics; whole-body tracking controllers execute robot references reliably but cannot replan an infeasible one. Recent language-to-humanoid systems bridge this gap by training. We measure how much of the gap closes with no training at all, by putting the deployment controller itself in the loop. Sample-simulate-select (S$^3$) draws $N$ motions per prompt from a frozen text-to-motion model, retargets each to a Unitree G1 by direction-matching inverse kinematics, rolls all of them out under full rigid-body dynamics with the pretrained SONIC tracking policy, and keeps the candidate the policy executed best. Because the verifier is the deterministic simulator itself, S$^3$ attains the any-of-$N$ ceiling by construction; what we measure is where that ceiling lies and what falls short of it. On 200 stratified HumanML3D test prompts with $N=8$, upright execution rises from 83.5% to 89.5% and hardware-gate passes from 33 to 85; on the complete test split (4,184 prompts) it rises from 80.5% to 89.5%. A kinematic verifier that predicts falls well (AUROC 0.90) recovers only a quarter of this gain: ranking a prompt's own candidates is harder than classifying the population. What selection cannot fix is one class, prompts that lower the pelvis, which a generator trained on retargeted robot data does execute. We further score the semantic fidelity of the executed motion with the standard text-motion evaluator, with a real-mocap control that attributes the loss to the robot projection, ablate the retargeter against GMR (complementary failures: the any-of-8 ceiling rises to 95.0% over both), and execute all 177 gate-selected clips on the real G1: every one completes standing, with hardware tracking error matching simulation ($r=0.94$).

Tue 22 SeptRoboticsComputer Vision and Pattern Recognition
The gist
Generating human-like robot motions from text is hard because robots have real physical limits that typical models don’t consider. The authors propose a method called sample-simulate-select (S³) that tries out multiple motions from an existing text-to-motion model, simulates each on a robot, and picks the best one that the robot can actually perform. This improves how often robots can successfully follow text commands without needing to retrain the models. Their experiments show noticeably better success rates on a wide range of test prompts.
Open → 2609.26420v1

Vision language action models shrink drastically with fast offline recovery

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

Abstract: Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%. A recent approach restores such a model with supervised fine-tuning followed by reinforcement learning, which needs online rollouts and hundreds of GPU-hours. We recover most of the lost success entirely offline. Width pruning narrows the blocks but keeps the residual stream at its original size, so teacher and student hidden states have the same shape and are matched directly, without a projector. Training against a cache built in one teacher pass lifts the 63%-reduced student to within 3.5 points of the teacher in about 8 GPU-hours. A sweep over nine ratios locates where the recovery objective starts to matter. Up to 45% reduction the two do not differ significantly on OpenVLA-OFT. Hidden-state distillation then adds +2.1 to +4.5 points there between 63% and 87%, and +9.4 to +22.1 points on CogACT from 63% onward. At 81% on CogACT, a tripled recovery budget narrows the distilled student's gap to the teacher to 3.9 points on average, while supervised recovery stays more than 20 points below. At matched compression, width pruning yields higher success and depth pruning lower latency. On a 6-DoF manipulator, the distilled student at 72% reduction reaches 77.5% success against 59.5% for supervised recovery, runs 2.23x faster on-board than the teacher, and uses 62% less memory.

Thu 17 SeptRobotics
The gist
Robots use big AI models that understand vision and language to follow instructions, but these models are often too large to run on actual robots efficiently. The authors show how to shrink these models by removing many parts and then quickly restore most of their capability using a method that only needs offline data. Their approach narrows the model’s size without losing much success in tasks, running faster and using less memory than before. This makes it easier to put powerful robot AI on real hardware without lengthy retraining.
Open → 2609.19579v1

FluxVLA Engine simplifies building robots that see and act

FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

Abstract: Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform that turns heterogeneous embodied-policy components into a reproducible data-to-deployment workflow. Rather than introducing another policy model, $\mathrm{FluxVLA}$ standardizes interfaces for datasets, visual-language and world models, action heads, reward- or advantage-weighted learning, distributed training, simulation evaluation, optimized inference, and robot operators. The engine further integrates compositional dual-arm simulation, scalable automatic data generation, and model-decoupled human-in-the-loop rollout, takeover, correction collection, and reward annotation. For responsive physical execution, it combines Real-Time Chunking (RTC) with accelerated inference backends, lightweight remote GPU serving, and configurable trajectory post-processing. Together, these capabilities connect offline learning, simulation validation, online correction, and real-robot execution through shared and auditable contracts. $\mathrm{FluxVLA}$ therefore targets the engineering bottlenecks separating promising embodied-learning algorithms from reproducible evaluation and dependable deployment. Code is available at https://github.com/FluxVLA/FluxVLA

Tue 15 SeptRoboticsArtificial Intelligence
The gist
Building robots that understand what they see and act accordingly involves many complex parts like different data types, training systems, and ways to test. The authors created the FluxVLA Engine, which is a platform that connects all these pieces in one place to make it easier to develop and use robot policies. Instead of making a new robot model, FluxVLA standardizes how all components work together, from training to running on real robots. This helps turn new robot learning algorithms into practical, reliable systems faster.
Open → 2609.17210v1

Action chunking transformer encoder removal impact revisited in robot learning

The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer

Abstract: Action Chunking Transformers (ACT) are widely used to learn robot manipulation from demonstrations. Their conditional variational autoencoder includes an encoder meant to capture differences between demonstrations during training. The original ACT paper reported that encoder removal dropped the mean success rate from 35% to 2% on two simulated tasks with human demonstrations. We re-ran this ablation in the original code and checked whether the findings depend on the implementation or training data. The published drop does not reappear in our tests, although smaller gains or losses in success rate remain uncertain. To investigate the discrepancy, we varied training length and how checkpoints are selected for evaluation. Both can reverse which policy scores higher, but the published drop's cause remains unknown. Success rates alone leave open whether the encoder provides information that helps the policy reconstruct demonstrated actions. On the tested ACT benchmark, the sampled latent provides little reconstruction benefit at every tested nonzero weight of the penalty on latent information. At inference, ACT leaves this latent unused and sets it to zero. Skipping the encoder increases training throughput in both implementations we timed. We release code, evaluation tools and results so others can repeat the comparisons and test the encoder on other tasks.

Tue 15 SeptRoboticsMachine Learning
The gist
The paper re-examines a key claim about robot learning models called Action Chunking Transformers (ACT). The original claim said removing part of the model called the encoder made the robot much worse at completing tasks. The authors reran tests and found this big drop did not reappear, suggesting the encoder might not be as crucial as thought. They also found the encoder’s internal information is not really used during task execution, and skipping it speeds up training. They share their code so others can check and explore further.
Open → 2609.16745v1

Compact vision language action models cut parameters without losing skills

Dense to MoE Adaptation for Compact Vision Language Action Policies

Abstract: Vision language action (VLA) policies continue to grow in parameter count, making deployment on resource-constrained robot platforms difficult. The central goal is to reduce the number of LLM-side parameters retained in the deployed policy while preserving downstream task performance. Our approach, AdaDE, adapts selected dense feed forward blocks into mixture of experts (MoE) layers and derives expert retention masks from router statistics during fine tuning. The Dense2MoE conversion preserves the original dense FFN function at initialization, so expert deactivation can start without a separate recovery stage. Instead of using a fixed shutdown rule, expert masks are updated dynamically from router usage statistics, with staged training and expert protection to avoid early collapse. With 40% of the LLM parameters deactivated, AdaDE retains 95.1% average success in LIBERO and 42.0% average success across all 50 RobotWin2.0 tasks. These results suggest that dense to MoE adaptation with dynamic expert deactivation is a practical direction for reducing active VLA model size without severe performance loss.

Tue 15 SeptRobotics
The gist
Modern robot models that understand vision and language often have many parameters, making them hard to run on small robots. The authors found a way to turn parts of these models into special blocks called mixtures of experts, allowing many parameters to be turned off while keeping most task performance. This switching is done smartly during training, so robots can use smaller models without learning to recover from shutting off parts. Their approach keeps success rates high while reducing active model size, showing a practical way to deploy complex robot skills on limited hardware.
Open → 2609.16503v1