Papers for

home robotics teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Visual language models struggle tracking moved objects out of sight

Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos

Abstract: Real-world AI systems must reason about objects that are no longer visible: an AR assistant guiding a user back to an object used earlier, a household robot retrieving an item someone put away. This requires not just recalling where an object was last seen, but updating its state when it is moved and retaining that update once it leaves view. We refer to this as out-of-sight spatiotemporal reasoning. We introduce Beyond3D, the first VQA benchmark to isolate this ability in dynamic egocentric video: every query targets an object that has been relocated and has since left the field of view. We create our questions from HD-EPIC annotations, building a visibility track for each dynamic object from its 3D position, the camera pose, and the scene geometry to understand at each moment whether it is visible, occluded, or out of view. Beyond3D comprises 9,000 questions in eight types over 135 videos from nine participants, organized as one reasoning chain: visual grounding (is the target observable now), temporal grounding (when it was last visible and last placed), scene localization (which fixture anchors that location), and 3D spatial perception (where it lies relative to the current viewpoint or another object in the scene). We benchmark nine general-purpose and spatially specialized VLMs. The best model reaches 42.2% against 29.7% chance and text-only baselines reaching 31.9%, with the largest failures in recovering when an object was last visible, showing that tracking object movement out of sight remains far from solved for current VLMs.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Knowing where an object is after it disappears from view is hard for AI, like helping a robot find something you put away earlier. The authors made a new test called Beyond3D to check if AI models can recall objects moved out of sight in videos filmed from a first-person view. They tested nine different models and found even the best one only got about 42% right, struggling especially to remember when the object was last seen. This shows current AI still has trouble tracking objects when they go out of the camera's view.
Open → 2609.34630v1