Papers for

robotics system engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Visual problem solving benchmark improves generative model planning

SolveEdit: Benchmarking Visual Problem Solving in Generative Models

Abstract: Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Many real-world problems involve changing images to reach a goal, like moving objects without messing up the rest of the picture. The authors created a test called SolveEpIT to check if computer programs can do these tasks well. They also made a new two-step method that helps programs plan changes before making them, which works better than previous ways. Even the best current program only solved about half the problems correctly, showing this is a hard challenge.
Open → 2609.35504v1

Communication aware inference method cuts compute and latency in connected robotics

ComVLA: Communication-Aware Split Inference for VLA Models in 6G-Connected Robotics

Abstract: Connected robotics is an emerging 6G application where mobile robots follow natural-language instructions to manipulate physical objects. The Vision-Language-Action (VLA) models that enable this are too large to run on the robot; a common trend is to offload inference to the cloud. The wireless link, however, limits how much sensing data the edge can transmit per control step. Two recent lines address this constraint: semantic communication codecs compress sensor data but require channel-specific retraining, and VLA token pruners select tokens from image but ignore the channel. Our insight is that the dense semantic information contained in the language already indicates which visual tokens matter. We propose ComVLA, a framework that uses this language guidance to adapt the VLA token budget to the channel capacity. Transmitting 32 tokens instead of 512 on the LIBERO benchmark, ComVLA cuts inference compute by 74% and inference latency by 22% versus the original OpenVLA-OFT baseline, at a cost of 1.5 pp in average task success (95.4% vs. 96.9%), and it stays within the capacity budget under Rayleigh and Rician fading. These results demonstrate that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.

Mon 7 SeptRobotics
The gist
Robots that follow spoken or written instructions need large AI models to understand and act. These models are often too big to run on the robot itself and must rely on a wireless connection to a powerful server. The authors found a way to reduce the amount of data sent over the wireless link by using the language part of the instruction to guess which parts of the image are important. This reduces computing work and delays while keeping the robot's success rate high, even with imperfect wireless signals.
Open → 2609.07838v1