Papers for

mobile robot developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

UAV memories combined on ground to improve low-altitude answers

Memory in the Sky: Low-Altitude Question Answering with Multi-Agent Memory Aggregation

Abstract: This paper studies low-altitude question answering (LAQA), in which distributed unmanned aerial vehicle (UAV) memories are aggregated at a ground server to answer questions about observations over a long horizon. Unlike conventional resource allocation based on sensing, communication, control, or computation metrics, LAQA requires an explicit measure of memory value. We propose a generative adversarial exam (GAE) that uses forward simulation to evaluate memory retrieval and exam scores to quantify memory quality. This enables the downstream QA value of candidate memories to be measured and optimized without accessing the internal mechanisms of the black-box captioning, retrieval, and reasoning pipeline. Building on this metric, we develop a memory-centric (MemCen) framework that jointly selects UAVs and allocates transmit power to maximize memory quality under communication constraints. In the noise-limited regime, we derive a QoM-aware capped water-filling law that explicitly connects task utility with physical-layer power allocation. We further develop penalty successive optimization (PSO) and learning to memorize (L2M) solvers. MemCen achieves QA accuracies of 92.4% and 84.0% in CARLA Town04 and Town05 under static and dynamic communication conditions, respectively. In real-world experiments, MemCen achieves 88.5% QA accuracy on the panoramic multi-agent system (PMAS) benchmark. Finally, UAV-to-robot-dog demonstrations further validate the practical utility of the acquired memories for environmental understanding and navigation.

Mon 28 SeptRoboticsInformation Theory
The gist
Answering questions about what drones see over time is hard because their memories are spread out and communication can be limited. The authors propose a way to score and pick the most useful drone memories to send to a ground server, even without looking inside complex AI systems. They also develop methods to pick which drones to use and how much power to send data using new formulas. Their approach works well in both simulated and real-world tests, helping robots understand and navigate environments better.
Open → 2609.35431v1

Panoramic vision improves robot navigation with longer planning

PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation

Abstract: Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward: more complete visual context should enable better-informed navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration. However, we find that simply replacing perspective images with panoramas yields only limited gains. Our diagnosis suggests that fully exploiting wider visibility requires modifications to action prediction, training supervision, and visual representation. First, wider visibility supports longer-horizon action planning. We make the model predict longer action sequences, enabling larger turns and subsequent movement from a single panorama. Specifically, we introduce a confidence-guided execution (CGE) strategy that dynamically determines how many predicted actions to execute before replanning. Second, wider visibility also brings more complex route choices. We therefore construct training routes with frequent branching points and clear instructions to provide targeted supervision for route selection. Third, panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks. We combine semantic and geometric features from RGB panoramas to capture both scene content and spatial layout without adding visual tokens. With a 4B backbone and RGB-only input, PanoVLN surpasses the previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen. Real-world experiments on a quadruped further demonstrate faster navigation with fewer pauses than prior VLN methods.

Mon 28 SeptComputer Vision and Pattern RecognitionRobotics
The gist
The authors studied how robots can better follow instructions in indoor spaces by looking around more completely using panoramic views instead of narrow camera shots. They found that just having a wider view isn't enough; the robot also needs to plan longer sequences of moves, be trained with routes that have more choices, and understand the spatial layout better. They built a method called PanoVLN that combines these ideas and tested it in various tasks, showing it helps robots navigate more successfully and efficiently. Their experiments included real-world runs on a robot dog, which moved faster and paused less than previous methods.
Open → 2609.34759v1

SparseNav improves indoor robot navigation with selective map labeling

SparseNav: Instruction-conditioned Sparse Semantic Perception for Training-Free Vision-Language Navigation

Abstract: Map-based vision-language navigation (VLN) relies on persistent spatial representations to connect language understanding with geometric planning. However, acquiring semantics beyond the needs of the current instruction can introduce unnecessary perception cost and irrelevant annotations. Continuously accumulating unrelated objects may not only waste computation, but also clutter the visual-spatial representation consumed by the vision-language model (VLM) planner. To address this problem, we present SparseNav, a training-free framework that follows a less-is-more principle for semantic navigation. SparseNav persistently maintains a lightweight geometric bird's-eye-view (BEV) map and sparse landmark memory, acquiring new semantics on demand using the active sub-instruction to decide what is worth grounding. An instruction manager first tracks navigation progress and identifies the active landmark query. An instruction-conditioned perception mechanism then invokes open-vocabulary segmentation when the queried landmark is visible and its metric location can inform the next decision. The resulting landmark memory supports VLM selection among hybrid frontier and local directional waypoint candidates. Without any additional training, SparseNav achieves success rates of 42.8% on R2R-CE and 40.7% on RxR-CE, both on the Val-Unseen splits. Controlled ablations examine semantic perception strategies and the contributions of individual framework components. Furthermore, we successfully deployed SparseNav on a Unitree Go2 quadruped equipped with an Intel RealSense D455 RGB-D camera for geometric mapping and landmark grounding and a Livox MID-360 LiDAR for localization, without a prebuilt map. We validated its effectiveness across multiple indoor environments using instruction-conditioned waypoint navigation.

Tue 22 SeptRobotics
The gist
Indoor robots need to understand both language instructions and their surroundings to find their way. The authors created SparseNav, a system that only adds important landmarks to a map based on the current instruction, saving computing power and avoiding confusion. It works without extra training and was tested successfully in simulated environments and on a real robot navigating indoors. This selective approach helps robots follow complex directions more efficiently.
Open → 2609.26408v1

Omnidirectional walking robot improves safety and energy use in obstacle tasks

Safety-Constrained Model Predictive Control for an Omnidirectional Walking Assistive Robot Using Control Barrier Function

Abstract: Providing safe and effective mobility assistance plays a crucial role in restoring independence and enhancing the quality of life for individuals with motor impairments. In this context, robotic walking assistive devices have recently emerged as promising solutions to provide physically compliant interaction while ensuring user safety and support. This paper presents a novel control framework for an omnidirectional Walking Assistive Robot (I-WANDER) that integrates a Control Barrier Function (CBF) formulation into a Model Predictive Control (MPC) scheme to explicitly enforce collision-avoidance safety constraints while optimizing for energy efficiency and smooth human-robot collaboration. The method was experimentally evaluated with 12 healthy participants performing two different walking tasks using both the proposed CBF-based MPC controller (CB-MPC) and a variable admittance controller (AC). The first task involved structured navigation through a U-shaped corridor, whereas the second consisted of a single-obstacle avoidance task performed blindfolded to ensure the obstacle was unexpected. Comparative results show that the CB-MPC architecture significantly reduces energy consumption and mechanical work (p < 0.01) without compromising motion smoothness, while also decreasing the number of obstacle collisions. Overall, the findings highlight the potential of the proposed control architecture to enhance both safety and efficiency in robotic walking assistance.

Tue 22 SeptRobotics
The gist
Helping people with motor difficulties move safely and comfortably is important. The authors designed a robot that helps users walk in any direction, using smart controls to avoid bumping into things and to save energy. They tested their new control system with people walking through tricky paths and avoiding unexpected objects. Results showed this method lowered energy use and collisions while keeping smooth movement, suggesting a safer, more efficient walking aid.
Open → 2609.25994v1

Foundation model confidence and depth improve detection of unknown objects

Combining Foundation Model Confidence and Monocular Depth for Training-Free Out-of-Distribution Segmentation

Abstract: Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo. The reliable detection and segmentation of these out-of-distribution (OOD) objects is therefore crucial for a safe understanding of the environment and decision-making. Most existing approaches require access to OOD training samples, retraining of the segmentation backbone, or dedicated auxiliary architectures, limiting their practical applicability. We propose a training-free method that derives dense OOD scores directly from the confidence predictions of a foundation segmentation model, without any task-specific fine-tuning or access to anomalous data. To improve the robustness of our OOD segmentation, geometric information from monocular depth estimation is incorporated into the decision process, providing complementary cues to uncertainty-based predictions. We evaluate the proposed method on the SegmentMeIfYouCan benchmark and additionally assess its performance on OOD tracking in video sequences, reflecting the temporal nature of real-world perception systems. The method performs strongly on road-centered benchmarks.

Sat 19 SeptComputer Vision and Pattern Recognition
The gist
Self-driving cars need to notice when they see things they haven't seen before, like strange animals or loose items on the road, to stay safe. The authors propose a way to detect these unusual objects without needing extra training or special data. They combine how confident a big AI model is about what it sees with depth information from a single camera to spot these unknown things more reliably. Their method works well on road-focused tests and in video sequences, helping cars better understand their surroundings.
Open → 2609.22896v1

Robot navigation improves with smart waiting and rerouting decisions

SPARROW: Survival-POMCP for Adaptive Robot Routing, Observation, and Waiting

Abstract: Temporary obstacles that may block a robot's planned route create a sequential navigation problem: a robot must decide whether to wait for a blockage to clear, reroute, or acquire more information about the obstacle before acting. We formulate graph navigation among temporary obstacles as a partially observable semi-Markov decision process and introduce SPARROW, a belief-space planner built on Partially Observable Monte Carlo Planning (POMCP). SPARROW searches over traversal, observation, and finite-duration waiting actions while maintaining a particle belief over latent obstacle classes and clearance times. Class-conditioned survival models are learned online from both clearance observations and right-censored encounters where the robot reroutes before clearance is observed. A generative model simulates obstacle arrivals and clearances as each action unfolds, so the planner can account for blockages that may occur along alternative routes. We further introduce a value-of-learning criterion that trades the immediate cost of collecting labelled survival data against its expected reduction in future navigation regret. Across two simulation graphs and multiple obstacle-class settings, SPARROW reduces mean time-to-goal by 12-26% relative to OSCAR, a recent survival-based method for the same problem. On a physical mobile robot, SPARROW reduces mean time-to-goal by 20.5% relative to OSCAR while selectively observing, waiting, and rerouting as environment conditions change.

Thu 17 SeptRobotics
The gist
Robots often face temporary obstacles that block their way, like a door that might open soon or a hallway that might clear. The authors created a method called SPARROW that helps robots decide whether to wait, find a new path, or gather more information before moving. SPARROW uses a technique that plans while considering uncertain obstacles and learns how long they usually last. This approach helps robots reach their destinations faster both in simulations and real-life tests.
Open → 2609.21008v1

Robots navigate safer using semantics in 3D uncertainty maps

SemSafe-3DGS: Semantic Risk-Aware Active Navigation in Uncertain 3D Gaussian Splatting Maps

Abstract: Autonomous robots operating in partially observed environments must navigate safely while acquiring observations that improve future planning. Existing safety formulations generally reason primarily about geometry. Consequently, geometrically similar scene elements may induce comparable control responses despite having different semantic consequences. We present a semantic risk aware safe-active perception framework for navigation in attributed 3D Gaussian maps. Semantic attributes modulate an Average Value-at-Risk collision clearance model through class dependent risk weights, allowing safety-critical Gaussian primitives to receive greater influence in the composite barrier. The resulting weighted clearances are aggregated into a control barrier function, while a trajectory-relevant active perception barrier promotes observations that reduce geometric map uncertainty along the robot's anticipated motion. Both objectives are integrated in a unified CBF-QP that enforces semantic risk-aware collision avoidance as a hard constraint while relaxing information acquisition when it conflicts with safety or task progress. Experiments demonstrate efficient safety constraint, improved navigation through active perception, semantic dependent trajectory adaptation, and real-robot execution under Ackermann dynamics.

Wed 16 SeptRobotics
The gist
Robots need to move around safely even when they don’t see their whole environment clearly. This work helps robots understand what different parts of a scene mean, not just how they look or where they are. By giving more importance to objects that might cause more serious problems if hit, the robot can plan safer paths and also decide where to look next to improve its map. The authors tested their system on real robots that drive like cars and showed it helps avoid risky situations better while exploring.
Open → 2609.19330v1

Robot plans take humans intentions into account using vision language models

HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models

Abstract: Approaches to incorporating human awareness into mobile robot decision-making mainly focus on collision avoidance in low-level motion planning, often overlooking the challenges posed by human presence and high-level behavior. To address this vacancy, we present HINT-Plan, a novel approach to integrate human intention prediction into robot task planning. HINT-Plan employs Vision Language Models (VLMs) to anticipate high-level human intentions from third-person image observations, convert them into goal states, and solve joint task-planning problems. To effectively enable scene awareness in context-rich environments, we use hierarchical Scene Graphs (SGs) as high-level representations of the environment, and translate environmental topology and actionable knowledge into formal planning language to ensure executable plans. Evaluated in a photorealistic simulation, HINT-Plan achieves an overall success rate of 69.71% in joint human-robot task planning, substantially outperforming the baselines by up to 35.29%, while also reducing functional conflicts. The results show the effectiveness of explicitly incorporating inferred human intentions into formal multi-agent task planning for proactive human-aware robot decision-making.

Tue 15 SeptRoboticsArtificial Intelligence
The gist
Robots often avoid bumping into people but don't usually understand what people want to do around them. The authors developed HINT-Plan, which helps robots predict what a person nearby intends to do by looking at images and using special language-vision tools. The robot then plans its own tasks considering these human intentions and the environment details. This method was tested in a realistic simulation and showed better success in working alongside humans than previous approaches.
Open → 2609.17771v1

Neural control barrier functions enable safer robot navigation

VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search

Abstract: As the number of autonomous robots continues to grow, safety becomes increasingly important. Control barrier functions (CBFs) provide a theoretically grounded framework for ensuring safety, but existing design methods often face limitations in effectiveness, scalability, or interpretability, and may result in overly conservative safe sets. In this paper, we propose \emph{VertexCBF}, a framework for learning neural CBFs in a scalable, systematic, and explainable way. We approximate the stationary Hamilton--Jacobi value function using a neural network trained via a combination of physics-informed and sparsely supervised learning. By exploiting control-affine dynamics and a convex polytope control set, under which the Hamiltonian is maximized at the control vertices, we efficiently generate supervision points via GPU-parallel vertex-restricted tree search, while a residual architecture guarantees that the learned CBF is never larger than the specified constraint function. We evaluate the method on 15 systems and compare it against relevant baselines, showing that it reliably recovers large safe sets where the baselines are conservative or fail completely. In addition, we perform a hardware experiment in which a mobile robot safely avoids pedestrians using a neural CBF trained with our method.

Fri 11 SeptRoboticsMachine Learning
The gist
Keeping robots safe is important as they become more common. Control barrier functions (CBFs) help robots avoid danger, but existing methods can be too cautious or hard to use. The authors created VertexCBF, a new approach that uses neural networks to better learn safe behaviors efficiently and understandably. Their method works on different robots, making safe zones larger and avoiding failures seen in older methods. They tested it on a real robot that safely avoided people walking nearby.
Open → 2609.12831v1

Pedestrian trajectory prediction improves using maps and awareness states

MamMA: A Mamba-Based Pedestrian Trajectory Prediction Algorithm Considering Occupancy Map and Pedestrian Awareness States

Abstract: Many pedestrian trajectory prediction algorithms have been proposed to improve the safety of navigation for mobile robots working in human-robot coexistence environments. Some pedestrian trajectory prediction algorithms extract information about obstacles near pedestrians from top-down view images to improve the accuracy of trajectory prediction. However, mobile robots typically create local occupancy maps using LiDAR, rather than top-down view images. Meanwhile, the vision sensors on board robots provide egocentric view images, which contain fine-grained behavioral information about the pedestrians near the robot. To better use the information collected by LiDAR and on-board vision sensors, we propose MamMA, a Mamba-based pedestrian trajectory prediction algorithm considering occupancy maps and pedestrian awareness states. MamMA divides the occupancy map by patches and extracts obstacle features from each patch to create map features. Pedestrian awareness states are divided and considered, as some studies show that awareness states affect the perception and speed of pedestrians. Furthermore, a Mamba-based model is proposed to predict the future trajectories of pedestrians based on different types of features. Experiments on the STCrowd, SiT, JRDB, ETH, and UCY datasets show that MamMA achieves better average displacement error and final displacement error than the state-of-the-art algorithms.

Mon 7 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Predicting where people will walk helps robots move safely around them. The authors developed MamMA, a method that uses maps from laser sensors and also considers whether pedestrians are aware of their surroundings, which affects how they move. They break down the map into parts to better understand obstacles and use a Mamba-based model to predict pedestrian paths. Tests on several common datasets show that MamMA predicts movements more accurately than prior methods.
Open → 2609.08041v1

Robot improves person following using maps to find lost targets

Human-Aware Target Tracking and Navigation: Fusing Kinematic State Estimation with Structural Map Constraints

Abstract: Autonomous mobile robots performing person-following tasks often suffer from temporary occlusions and sensor track loss in dynamic environments. This research presents an end-to-end autonomous navigation stack that addresses target occlusion through map-informed spatial reasoning. The proposed system features a multi-modal perception pipeline, fusing deep learning-based visual tracking with 2-dimensional LiDAR point clustering to maintain high-fidelity tracking of a tagged person. A continuous state estimator integrates this perception data with wheel odometry and IMU sensors for stable localization. When the active track is lost due to occlusion, the system activates a map-based recovery framework. Leveraging a predefined topological map, the system executes a graph-based search to propagate the target's last known trajectory along structurally defined walking lanes, adhering to left-hand regional conventions. By generating a discrete set of feasible future trajectories, the robot reasons about potential structural trajectory changes, such as continuing a heading or turning at an intersection. This map-informed prediction is fed directly to the local obstacle avoidance planner, enabling the robot to continue following its target safely and predictably until the person is visually reacquired. Real-world evaluations in dense multi-person environments demonstrate the system's robustness, achieving a 71.4\% target reacquisition success rate during major occlusion events lasting up to 7 seconds.

Mon 7 SeptRobotics
The gist
Robots that follow people often lose track when the person is hidden behind obstacles like other people or walls. The authors created a system that helps the robot keep following by combining video and laser sensors with a map of the environment. When the robot loses sight, it uses the map to guess where the person might be and moves intelligently to find them again. This approach improved the robot’s success in re-finding people during real-world tests.
Open → 2609.07091v1