Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

2026-08-24Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial Intelligence
AI summary

The authors developed a model called Object-Uni to better understand and control the exact positions and orientations of objects in images. Unlike previous systems that only describe objects, their model treats an object's pose as a core piece of information that links how it is seen and how new images are created from different viewpoints. They introduced a new way to describe object orientation that large language models can use, and created a large dataset to help train their model. Their experiments show improved ability to recognize object poses and generate images with accurate spatial consistency.

object posespatial understandingpose-conditioned generationnovel view synthesismultimodal modelsgeometric supervisionlarge language modelsobject-centric representationviewpoint abstractionbenchmark dataset
Authors
Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen, DanDan Zheng, Libin Wang, Weiming Dong
Abstract
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.