Papers for

robotics integration teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

ForceTwin identifies precise physics for better robot object handling

ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction

Abstract: Manipulating objects requires understanding not only their motion, but also the physical properties that determine it. For articulated objects, these include inertia, friction, and mechanisms such as springs or door closers, whose effects can vary with configuration and velocity. Such properties are not directly observable from appearance: visually identical doors may require very different effort to manipulate. Existing digital-twin pipelines recover primarily kinematics or assign static physical parameters from visual and language priors, which can yield physically implausible estimates. As a result, state-dependent mechanism dynamics remain unidentified and are not represented in standard asset formats. We present ForceTwin, a system for identifying physics-informed digital twins of articulated objects from instrumented human interaction. A person probes an object using a handheld force-sensing gripper, providing synchronized poses and interaction forces from which we estimate the articulation, parametric dynamics including inertia, Coulomb friction, viscous damping, and a structured neural residual capturing state-dependent mechanism forces. ForceTwin nearly halves the inertial-parameter error of a VLM prior. As a feedforward dynamics model for impedance control on a Spot and a Franka FR3, ForceTwin achieves 87% goal completion across nine object-embodiment pairs, compared with 60% using VLM-prior and 57% using kinematics-only twins, with the largest gains on objects whose strong mechanisms cause both baselines to stall. We further use the identified twins to train whole-body door-traversal policies and deploy them in the real world. Project Page: https://timengelbracht.github.io/forcetwin-website/

Fri 18 SeptRoboticsArtificial Intelligence
The gist
Robots have trouble opening or moving objects like doors because they don’t know the forces involved, such as friction or spring effects. The authors created ForceTwin, which watches a person use a special tool to measure force and movement on these objects. This data helps build a detailed, physics-based model so robots can predict how objects will behave. Using these models, robots can manipulate objects more successfully, especially ones with tricky mechanisms.
Open → 2609.21751v1

UAV navigation aligns language commands with diverse safe flight paths

VLM-MPPI: Grounding Natural Language in Behaviorally Diverse Trajectories for Aerial Navigation

Abstract: We present a hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments. To bridge the gap between abstract semantics and low-level control, we employ a parallelized ensemble of six behavior-conditioned Model Predictive Path Integral (MPPI) planners. Crucially, by designing mode-specific guiding costs and sampling biases, we induce distinct trajectory modes that converge to unique behavioral means, yielding a compact set of intentionally diverse candidates rather than mere stochastic variations. We project these 3D candidates onto the onboard first-person-view RGB stream, turning language grounding into a visual action selection problem. A pretrained vision--language model (VLM) asynchronously selects the candidate index given the overlaid FPV image and a natural-language prompt, while MPPI replans at 20Hz and a PID-based low-level controller tracks the selected trajectory. We implement the full pipeline in NVIDIA Isaac Sim and on a real-world quadrotor platform equipped with LiDAR and RGB sensing. Experiments in both simulation and real-world flights show semantically meaningful behavior diversity, robust language alignment despite VLM latency, and safe, repeatable flight across all modes, achieving 100% task success in our evaluated scenarios.

Wed 16 SeptRobotics
The gist
Flying drones indoors is tricky because they need to avoid obstacles while following instructions. The authors created a system that plans six different safe flight paths based on what the drone is told to do. They turn the problem of understanding language into picking the best flight path by matching pictures the drone sees with the instruction using a visual-language AI model. Their system works well both in computer simulation and on a real drone, flying safely and successfully every time.
Open → 2609.18451v1

Large language models steady networked control systems with slow supervision

Large Language Models in the Loop: A Stability- and Network-Aware Survey in Networked Control, Cyber-Physical, and Multi-Agent Systems

Abstract: Modern networked control systems (NCSs), cyber-physical systems (CPSs), and complex multi-agent network systems (CNSs) increasingly rely on large language models (LLMs) for high-level decision-making. However, the slow, stochastic nature of LLMs directly conflicts with the strict stability and safety guarantees required by these physical systems. This survey presents a unified analysis of how LLMs can be admitted into the control loop of NCS, CPS, and CNS without compromising closed-loop guarantees. We organize this around a core principle: the LLM operates as a slow supervisor adjusting high-level goals and constraints, while a fast, certified inner loop maintains physical stability. Under this framework, LLM integration maps directly to classical networked control challenges, where inference latency acts as delay, API failures as packet dropouts, tokenization as quantization, and hallucinations as bounded disturbances. We assess current developments across all these three domains, highlighting that rising model capabilities are frequently accompanied by a drop in formal safety assurances. Finally, we propose concrete future research directions, identifying the widespread lack of formal stability proofs as the field's central open problem.

Tue 15 SeptArtificial Intelligence
The gist
Some computer systems that control physical devices or multiple agents need to be very reliable and stable. The paper's authors review how large language models (LLMs), which usually take time to respond and can be unpredictable, can still be included in these systems safely. They explain that if the LLM acts slowly to set high-level goals while a faster, certified system handles immediate control, stability can be maintained. They also compare typical LLM behaviors to known network control problems and say more research is needed to guarantee safety formally.
Open → 2609.16599v1