Surface based positional encoding improves uv texture generation for 3d models

DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding

Computer Vision and Pattern Recognition

Summary

Creating detailed and consistent textures for 3D models is hard because traditional methods struggle with hidden areas and mismatched seams. The authors developed DirectUV, a new technique that uses clues from the actual 3D surface instead of just the flat texture layout to guide texture creation. This makes the textures look sharper and more consistent, especially in parts usually missed or distorted by other methods. Their approach relies on teaching the system to understand where each texture bit belongs on the 3D shape, improving the overall quality.

What this means in practice

  • For 3d artists: Generate higher quality textures for 3D models from a single image while preserving surface details and seam consistency.
  • For game developers: Improve texture mapping on complex 3D characters or objects to fill occluded or unseen regions accurately, enhancing visual fidelity.

Authors

Jiantao Lin, Yingjie Xu, Mingzhi Sheng, Yangkai Wei, Hao Chen, Ying-Cong Chen

Abstract

Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid positions rather than to the underlying surface geometry. This mismatch limits coherence across seams and disconnected UV islands. We propose DirectUV, an image-conditioned UV texture diffusion framework that operates in the latent UV space of a pretrained image VAE, in which a Diffusion Transformer denoises the UV latent given a single input image and a coarse UV map. At its core, Surface-Aware Positional Encoding (SAPE) replaces the standard 2D-grid positional encoding with encodings derived from per-token 3D surface coordinates obtained via UV-to-surface correspondence. As positional encodings define the distance metric used by attention, SAPE enables tokens to interact according to 3D positional proximity derived from surface correspondence rather than UV-grid distance, restoring coherence across seams and disconnected islands. A multi-level extension further assigns different attention heads to progressively finer subdivisions of the same latent UV patch, allowing the model to reason about surface structure at multiple granularities. Experiments show that DirectUV produces sharper and more globally consistent textures than other baselines, with the largest improvements in occluded and view-unseen regions where projection-based methods leave gaps or stretched textures.