Papers for

robotic perception teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Vision language models assessed for true visual causal reasoning skills

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

Abstract: Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 15 multimodal models show that constraint sensitivity is task- and model-dependent: intervention has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRV-Bench-Constraint-Based-Evaluation-of-Causal-Reasoning-in-Vision-Language-Models

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Many AI systems that combine seeing and understanding language appear to do well on tasks, but it’s unclear if they really understand cause and effect from images or just guess based on patterns they’ve seen before. The authors created a special test called CCRV-Bench to better check if these models truly reason about causes in single images by controlling tricky clues that can mislead AI. They tested 15 models and found that some tasks and constraints reveal weaknesses that overall scores hide, showing that good scores don’t always mean real understanding. Their benchmark helps pinpoint exactly where these models succeed or fail in causal thinking grounded in what they see.
Open → 2609.30979v1

Multimodal fusion improves robotic new view synthesis from camera and LiDAR

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

Abstract: Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.

Thu 24 SeptRoboticsComputer Vision and Pattern Recognition
The gist
Generating new images of a scene from different viewpoints is a key problem for robots, but existing methods typically use only camera images. The authors show that combining camera images and LiDAR data, which measures 3D points directly, improves the quality of new view images and depth maps. Their method, called M3GD, uses pretrained models for images and point clouds and cleverly aligns their features without extra training. This improves the robot’s understanding of 3D space and lets it better synthesize novel views in real-world settings.
Open → 2609.30056v1