Papers for

augmented reality designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Embodied agent benchmark tests active visual reasoning in real scenes

JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments

Abstract: In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active reasoning and evidence acquisition. We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning agentic method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning. Code and benchmark are available at https://github.com/ControlNet/JRDB-AVR.

Mon 28 SeptArtificial IntelligenceComputer Vision and Pattern RecognitionRobotics
The gist
Sometimes robots or AI agents need to look around actively to find the right clues to answer questions about their surroundings. The authors created a new test called JRDB-AVR that checks if these agents can ask to see certain views or moments before giving an answer. They found many AI systems give correct answers without actually seeing the evidence, which means they guess or use shortcuts. Their benchmark helps measure both the correctness of answers and whether agents truly see and use the right information.
Open → 2609.35032v1

XR pen improves robot control speed and precision in home tasks

Comparative Evaluation of an XR Pen-based Control Interface for Semi-Autonomous Mobile Robot Navigation in Service Environments

Abstract: Service robots remain difficult to deploy in domestic environments, partly because fully autonomous operation is not yet reliable in unpredictable surroundings, and partly because conventional control methods remain inaccessible to novice users. Extended Reality (XR) enables operators to visualize robot information overlaid onto the real world and to interact with augmented elements. Yet, common XR control methods, such as motion controllers and hand gestures, are still perceived as unintuitive. This paper presents a control interface that uses a commercial XR pen to command a semi-autonomous mobile robot in Augmented Reality (AR): the operator points at a position in the room, selects it, and drags an augmented arrow to set the desired orientation of the robot at this destination. Two additional interfaces, based on the XR motion controllers and hand gestures, were developed within the same framework. To assess the performance and users' perception of these interfaces, and of the XR pen in particular, a study with 10 participants compared four control methods, i.e., the XR pen, the XR motion controllers, hand gestures, and a computer-based baseline RViz, in navigation tasks performed in a home-like environment. Results show that the XR pen significantly outperforms the other methods in task selection time with the most consistent selections, and that the XR motion controllers obtain the best perceived workload and usability scores, ahead of the computer-based baseline, supporting XR-based control as an intuitive alternative for novice users. However, technical limitations in the integration of the recently released XR pen currently hold back its user experience.

Fri 25 SeptRoboticsHuman-Computer Interaction
The gist
Controlling robots at home can be tricky because they don’t work well on their own and controlling them is often hard for beginners. The authors tested a new way to guide a robot using an augmented reality pen that lets a user point and drag arrows to tell the robot where to go and how to face. They compared this pen to other XR controllers, hand gestures, and a computer method in a home-like setting. The pen made picking tasks faster and more accurate, while the XR controllers felt easiest to use, though some technical issues with the pen remain.
Open → 2609.31117v1

Mind2Cloud creates 3D shapes from brainwaves with finer detail

Mind2Cloud: EEG-to-Point Cloud Generation with Two-Granularity Diffusion Decoding

Abstract: Reconstructing 3D objects from brain signals offers a promising avenue for understanding human visual cognition. While prior work has shown initial success using EEG signals for 3D reconstruction, existing methods typically employ a uniform diffusion decoder, overlooking the evolving semantic granularity of both EEG representations and the diffusion denoising process. In this paper, we propose Mind2Cloud, a novel EEG-to-point-cloud generation framework based on two-granularity diffusion decoding. The core of Mind2Cloud is a time-aware decoder that integrates a global Transformer branch and a local Point-Voxel CNN (PVCNN) branch across diffusion timesteps through a learnable fusion mask. Specifically, Transformer layers are incorporated into the early upsampling stages to capture global object structure under high uncertainty, while PVCNN modules are used in later stages to refine local geometric details. Inspired by the hierarchical nature of EEG-based visual representations, this design dynamically adapts its spatial granularity in accordance with the coarse-to-fine trajectory of diffusion denoising. We further introduce an adversarial refinement module to enhance geometric realism and semantic consistency. Extensive experiments on the EEG-3D dataset across all 12 subjects demonstrate that Mind2Cloud outperforms prior work in both geometric accuracy and semantic alignment, setting a new benchmark for EEG-to-point-cloud generation. Our source code is available at https://github.com/duasoi/Mind2Cloud.

Sat 12 SeptComputer Vision and Pattern Recognition
The gist
Turning brain signals into 3D shapes can help us understand how people see objects. The authors developed Mind2Cloud, a method that uses brainwave data (EEG) to rebuild 3D objects as point clouds by gradually improving from rough shapes to detailed ones. Their technique uses two types of neural networks working together to capture broad shapes first and then refine details. Testing on brain data from 12 people showed this method makes more accurate and realistic 3D shapes than previous attempts.
Open → 2609.13991v1

Image based 3D objects generated more accurately with frequency based method

Single Image to Textured 3D Object Generation in Frequency Domain: From Theory to Pipeline

Abstract: Single-view 3D reconstruction, also known as image-to-3D, is a persistently challenging task due to the extreme lack of information. Recently, diffusion models pre-trained on large-scale datasets served as 2D priors are used to solve the ill-posed task but suffer from color deviation and view inconsistency, which can be curbed by using diffusion models fine-tuned with 3D annotated data served as 3D priors. However, 3D priors lack high-frequency details, which cannot be solved by direct complementation with 2D priors in spatial domain for introducing erroneous low-frequency 2D prior guidance. In this paper, we revisit the characteristics of different diffusion priors from the frequency perspective. Based on our observations, we theoretically present a unified framework of hybrid optimization using multiple diffusion priors in frequency domain. Under this framework, we further propose Morpheus3D, a pipeline of 3D object generation from any single unposed image in the wild. Morpheus3D enhances 3D prior with high-pass image-prompt 2D prior guidance to reconstruct high-quality 3D objects while effectively suppressing view inconsistency, low-frequency color deviation, and high-frequency lacking problems. Both quantitative and qualitative experiments on the public and our collected datasets with complex textures show that our method exhibits significant improvements in generation quality.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
Turning a single photo into a detailed 3D model is very hard because one image has limited information. The authors looked at how existing AI tools handle details and colors differently in various parts of the image's frequency spectrum. They designed a new approach that mixes 2D and 3D AI models more cleverly by focusing on different frequency parts, which helps create 3D objects with better details and consistent colors from just one unposed image. Their new pipeline, Morpheus3D, shows better results on various textured objects than earlier methods.
Open → 2609.07085v1