Papers for

robotics simulation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Humanoid robots walk safer using model-informed reinforcement learning

Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction

Abstract: Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.

Mon 28 SeptRobotics
The gist
Humanoid robots need to move safely in complex environments, but traditional methods can struggle with unpredictability. The authors combined a simplified physics model with modern reinforcement learning to teach robots safer walking steps. They created a safety check that adjusts robot movements in real time to prevent falls. Their approach reduced unsafe events in tests with a simulated humanoid robot, though it sometimes caused bigger side-to-side movements.
Open → 2609.34486v1

Driving world model system generates realistic multi-view and LiDAR simulation

HelloWorld: Towards Practical Applications of Generative Driving World Models

Abstract: Driving world models provide a promising route toward scalable counterfactual data generation and interactive simulation beyond recorded driving logs. Realizing this potential requires a system that can generalize across diverse scenes, respond faithfully to prescribed controls, generate coherent multi-sensor observations, and operate efficiently under repeated inference. We present \textbf{HelloWorld}, a 2B driving world model system designed around these requirements. HelloWorld progressively specializes broad visual and motion priors from heterogeneous video data into controllable driving generation using ego pose, HD maps, and 3D boxes. A block-causal generation interface, together with adaptation to self-generated context, aligns the model with sequential simulation. The system further supports synchronized seven-camera RGB generation and conditional LiDAR synthesis, and is distilled toward few-step inference for efficient deployment. Experiments evaluate visual quality, control fidelity, cross-view consistency, robustness under repeated generation, inference efficiency, and LiDAR synthesis. Together, HelloWorld provides a unified framework for scalable driving data generation and interactive simulation.

Thu 24 SeptComputer Vision and Pattern Recognition
The gist
Simulating driving scenes is hard because the system must create realistic views from many cameras and sensors while following control inputs over time. The authors built HelloWorld, a model that learns from lots of driving videos to generate scenes that match given car paths and map data. HelloWorld can produce images from seven cameras and LiDAR data that stay consistent over time, allowing for efficient and interactive simulation. This helps create more realistic virtual driving data beyond just replaying recorded videos.
Open → 2609.28931v1

Code based system builds persistent worlds with video visuals

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

Abstract: We introduce \textbf{CoDeR}, a new paradigm for world modeling. Unlike existing video world models that implicitly represent world dynamics through visual observations, our system explicitly constructs an executable world with code and employs video generation models for visual realization. Specifically, we coordinate five complementary roles to translate high-level concepts into structured world rules, executable dynamics, and perceptual observations. This design enables \textit{long-term memory}, \textit{open-ended interactions}, \textit{autonomous world evolution}, and \textit{multi-agent scenarios}, where multiple entities can act, interact, and evolve persistently beyond the current observation. Extensive experiments demonstrate that our framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings. Code and model weights will be made publicly available. Project Page: \href{https://becauseimbatman0.github.io/CoDeR}{CoDeR}.

Tue 22 SeptComputer Vision and Pattern Recognition
The gist
Many computer models that simulate the world only use images and videos to show what happens, but this new system called CoDeR actually writes and runs code to create a world that can keep evolving. The authors combine several parts that help turn simple ideas into detailed rules, actions, and visuals you can see. Because it uses code, the world can remember things for a long time and let multiple characters interact and change over time beyond what is immediately visible. Their experiments show it works better than older models in keeping complex, long-term scenes and interactions.
Open → 2609.26458v1

Sample simulate and select improves text to robot motion without training

Sample, Simulate, Select: Physics-in-the-Loop Text-to-Motion for Humanoids Without Training

Abstract: Text-to-motion models generate plausible human motion but do not model a robot's dynamics; whole-body tracking controllers execute robot references reliably but cannot replan an infeasible one. Recent language-to-humanoid systems bridge this gap by training. We measure how much of the gap closes with no training at all, by putting the deployment controller itself in the loop. Sample-simulate-select (S$^3$) draws $N$ motions per prompt from a frozen text-to-motion model, retargets each to a Unitree G1 by direction-matching inverse kinematics, rolls all of them out under full rigid-body dynamics with the pretrained SONIC tracking policy, and keeps the candidate the policy executed best. Because the verifier is the deterministic simulator itself, S$^3$ attains the any-of-$N$ ceiling by construction; what we measure is where that ceiling lies and what falls short of it. On 200 stratified HumanML3D test prompts with $N=8$, upright execution rises from 83.5% to 89.5% and hardware-gate passes from 33 to 85; on the complete test split (4,184 prompts) it rises from 80.5% to 89.5%. A kinematic verifier that predicts falls well (AUROC 0.90) recovers only a quarter of this gain: ranking a prompt's own candidates is harder than classifying the population. What selection cannot fix is one class, prompts that lower the pelvis, which a generator trained on retargeted robot data does execute. We further score the semantic fidelity of the executed motion with the standard text-motion evaluator, with a real-mocap control that attributes the loss to the robot projection, ablate the retargeter against GMR (complementary failures: the any-of-8 ceiling rises to 95.0% over both), and execute all 177 gate-selected clips on the real G1: every one completes standing, with hardware tracking error matching simulation ($r=0.94$).

Tue 22 SeptRoboticsComputer Vision and Pattern Recognition
The gist
Generating human-like robot motions from text is hard because robots have real physical limits that typical models don’t consider. The authors propose a method called sample-simulate-select (S³) that tries out multiple motions from an existing text-to-motion model, simulates each on a robot, and picks the best one that the robot can actually perform. This improves how often robots can successfully follow text commands without needing to retrain the models. Their experiments show noticeably better success rates on a wide range of test prompts.
Open → 2609.26420v1

Video diffusion models’ attention flaws cause physics errors in motion generation

Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

Abstract: Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While existing solutions rely on external priors or specialized data, we investigate the root cause by exploring the internal mechanisms of these models. Specifically, we present the first interpretability study on the ''motion planning'' process of text-to-video diffusion models, revealing how motion trajectories form during early denoising stages. Building upon the ''first shape, then details'' finding, we combine cross-attention trajectory patterns with causal head contributions to identify a specific subset of attention heads driving motion planning. Further, our self-attention analysis shows that Rotary Position Embedding (RoPE) induces excessive spatial attention decay. This causes early candidate regions to prematurely lock into physically implausible positions, suppressing reasonable trajectories in adjacent frames and triggering generation failure modes. To address this fundamental flaw, we propose a lightweight architectural modification that scales the frequency of RoPE across different denoising steps. This strategy reduces excessive attention decay, helping the model explore better candidate regions to establish coherent physical motion. Finally, training-free and training-based experiments confirm the effectiveness of our approach in enhancing the physical commonsense of generated videos.

Sun 20 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Video diffusion models can create impressive videos but often make mistakes that break real-world physics, like impossible movement. The authors studied how these models plan motion during early stages and found that a part called Rotary Position Embedding (RoPE) causes the model to fix on wrong positions too early. This stops the model from creating realistic movement paths across frames. They propose a simple fix to adjust how RoPE works during video creation, which helps the model produce more physically believable motions without extra training.
Open → 2609.23658v1

LYRIC enables realistic whole-body object interaction from language commands

LYRIC: Language-Driven Physics-Based Character Control for Contact-Rich Whole-Body Object Interaction

Abstract: We present LYRIC, a generative flow-matching controller for language-driven physics-based contact-rich interaction control, that enables simulated characters to perform contact-rich whole-body object interactions from a free-form language instruction and a sparse terminal object goal. To obtain reliable expert trajectories from imperfect motion-capture references, a single tracking policy is trained using geometry-conditioned interaction rewards and relaxed reference tracking near hand-object contact. To guide interaction progress without prescribing a full-body kinematic reference, we factorize the controller into a task-level planner that predicts short-horizon object and humanoid-root trajectories, and an action generator that resolves whole-body motion and contacts in closed loop. After behavior cloning, we freeze the planner and post-tune the action generator on policy using the planner's predictions as stable supervision for intermediate task progression. In a controlled OMOMO evaluation, our tracker achieves 64.3% success compared with 53.2% for an InterMimic reimplementation, while a unified policy achieves 76.5% on the full OMOMO dataset. On the held-out split, LYRIC achieves 90.3% task success, compared with 74.2% for the strongest matched kinematic-planner baseline, with better semantic alignment and motion quality. Without retraining, the controller also supports test-time object-waypoint guidance. Qualitative results further demonstrate robust, natural contact-rich interactions and zero-shot transfer to novel object shapes. The webpage is available at https://neu-vi.github.io/LYRIC/

Thu 17 SeptRoboticsGraphics
The gist
Controlling virtual characters to interact with objects in realistic ways is hard, especially when instructions are given in plain language. The authors developed LYRIC, a system that lets simulated characters perform complex whole-body tasks using only natural language instructions and a goal for the object they interact with. LYRIC uses a two-part controller to plan short moves and then generate detailed body motions and contacts. It performs better on standard tests than past methods and can handle new objects without extra training.
Open → 2609.19688v1

Large language models generate robot walking goals for better training

From LLM-Generated Specifications to Learned Quadruped Locomotion

Abstract: Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavily on hand-crafted reward functions. Designing reward functions requires substantial manual engineering, and it is often unclear which local rewards will induce the desired global behavior. Shaped rewards from formal specifications in languages like Signal Temporal Logic (STL) can make rewards more interpretable, but writing STL specifications itself still requires domain expertise. We study whether large language models (LLMs) can fill this gap by generating Parametric Signal Temporal Logic (PSTL) specifications that are subsequently used for policy learning. Given a natural language locomotion objective and a constrained specification grammar, GPT-5.5 and Qwen 3.6 independently propose STL templates for command tracking, safety, and gait structure. We instantiate the parameters of the generated PSTL templates using expert trajectories and retain only specifications that are consistent with demonstrated expert behavior. The resulting specifications are then transformed into smooth, finite-history reward functions and used to train a quadruped locomotion policy with Proximal Policy Optimization (PPO) in MuJoCo XLA (MJX). We evaluate both \emph{gait-aware} and \emph{gait-agnostic} settings. The former specifies walking-trot, trot, and bound regimes, while the latter allows contact patterns to emerge from the task objective. We compare against hand-engineered rewards, Text2Reward-style LLM-generated reward code, and an expert-switching oracle. Gait-aware Qwen 3.6 specifications achieved 100\% survival and command success across all tested speeds (0.3--2.1 m/s) and matched the target gait at high speeds, whereas Text2Reward achieved 0\% for both metrics at $\geq 1.9$ m/s. Videos: https://stl-locomotion.github.io/

Mon 7 SeptRoboticsArtificial Intelligence
The gist
Training four-legged robots to walk properly usually needs experts to manually create reward rules, which is hard and unclear. The authors study if advanced language AI can write formal goal descriptions from simple instructions instead. They show that two language models create rules that match expert walking patterns, which are then used successfully to train robot walking policies in simulation. The AI-generated rules perform better than hand-made or code-generated rewards, especially for higher speeds and specific walking styles.
Open → 2609.07111v1