Papers for

virtual environment developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Visual symbolic agent improves language guided actions in virtual humans

A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans

Abstract: Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations), a visual-symbolic framework for language-driven vh that leverages a pretrained vlm with tool calling to unify perception, reasoning, and action within a single control loop. A.D.A.M.O. maintains a dual visual-symbolic world model that combines egocentric visual input and synchronized symbolic state to support grounded task-oriented behavior from natural language prompts. To support diagnostic evaluation, we introduce a controlled task suite organized by a cd taxonomy that breaks down spatial tasks into procedural and linguistic complexity. Experiments in controlled scenes show that semantic labeling strongly influences task completion and failure modes, reducing perceptual ambiguity while shifting failures toward downstream execution, whereas reasoning errors remain comparatively rare.

Mon 28 SeptArtificial IntelligenceGraphics
The gist
Virtual humans need to see, think, and act based on language instructions in a 3D world. This paper presents A.D.A.M.O., a system that combines visual input and symbolic understanding to help virtual humans perform tasks from natural language prompts. The authors show A.D.A.M.O. can better understand tasks by using labeled visual information, which reduces confusion but shifts some errors to action execution rather than reasoning. They also introduce a way to test these tasks with increasing complexity.
Open → 2609.35463v1

Simple gripper model helps robots fold cloth realistically

A Simple Gripper Interface for Simulator-Agnostic Cloth Manipulation

Abstract: This paper presents a grasping model for cloth manipulation specifically tailored to ease the deployment of robotic control methods. The model is robust, fast and easy to implement avoiding at the same time contact and friction considerations between the gripper and the cloth in favor of simple positional constraints. The gripper is described by its pose, jaw state, and an attached grasping volume. Two kinds of grasping volumes are considered: an axis-aligned box to simulate a pinch grasping and a square pyramidal volume to simulate point grasping. When the gripper closes, the discrete cloth positions lying inside this volume are selected, stored in the local gripper frame, and then transported with the gripper motion. A simple squeezing step is also included to progressively move the selected cloth positions toward the center of the grasping region, avoiding an instantaneous displacement at closure. The model can be used in any simulator as it only requires access to discrete cloth positions and a mechanism for imposing target positions as constraints. We implement our grasping model in conjunction with a constraint-based inextensible cloth simulator, where grasping is implemented as moving positional equality constraints coupled with stretch, shear, collision, and table contact projection steps. The same gripper trajectory is applied on a robot arm to fold a real piece of cloth, serving as a simple bridge between simulation and physical cloth manipulation and showcasing the realism and practicality of our idealized grasping model.

Thu 24 SeptRobotics
The gist
Robots have trouble grasping and folding cloth because cloth moves in complex ways. The authors created a simple model for robot grippers that picks up cloth by selecting cloth points inside a virtual grasping area without worrying about friction or contact details. This model works in lots of simulation systems and can also control a real robot to fold cloth, making it easier to test and deploy cloth manipulation methods.
Open → 2609.29340v1

GLAM model improves robot exploration and navigation in virtual spaces

GLAM: Training a latent world model over global spatiotemporal memory for active exploration and navigation

Abstract: Active exploration and semantic navigation require an embodied agent to build memory from partial observations, predict how the evolution of observed spatial memory may support future motion, and convert that prediction into actionable plans. We present GLAM, a goal-conditioned latent world model trained over global spatiotemporal memory, and GLAM NAV, the complete navigation system built around it. Given historical map tokens, a navigation goal, and the current robot pose, GLAM jointly predicts future map representations and robot-centric waypoint latents, allowing future spatial context and navigation intent to be inferred in a shared representation space. The model follows a JEPA-like latent prediction paradigm, operates directly on map-level latent tokens rather than RGB reconstruction, and uses a pretrained waypoint encoder-decoder to supervise and decode navigation plans within GLAM NAV. Training data are collected by replaying ObjectNav expert trajectories in Habitat over HM3D v0.2 scene assets and slicing them into multi-timescale prediction samples. On a controlled HM3D-ObjectNav subset reproduction setting, GLAM NAV improves over a reproduced BSC-Nav baseline in both success rate and success weighted by path length.

Sun 13 SeptRobotics
The gist
Robots need to remember what they've seen and predict what might happen next to explore and navigate effectively. The authors developed GLAM, a model that helps robots create and use a map-like memory to plan where to go next. This system can predict future maps and navigation steps jointly, so the robot can better understand both its surroundings and its goals. Tests in virtual environments showed that GLAM NAV, the full navigation system built with GLAM, was more successful at reaching targets than some previous methods.
Open → 2609.14561v1