Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

2026-08-24Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionMachine Learning
AI summary

The authors developed a new method to create very realistic pictures of people from different angles and at high resolution. Their approach improves how consistent and detailed the images look when viewed from multiple cameras at once. They trained their model on synthetic images of human faces with diverse looks and apparel, using a special training process that needs less data but still produces sharp, lifelike images. Additionally, they combined their system with another model to build accurate 3D face representations from many views. Their work shows that their technique is good at making clear, consistent multi-view images of people.

photorealistic synthesisnovel view synthesisnext-scale autoregressive modelcross-view consistencysynthetic datasetmulti-view outputstransformer3D Gaussian liftingpixel-aligned modelface reconstruction
Authors
Federico Stella, Fei Jiang, Zhongshi Jiang, Zohar Barzelay, Emanuel Garbin, Amin Jourabloo, Liuhao Ge
Abstract
Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving identity, fine appearance details, and geometric coherence is critical. We build on the next-scale autoregressive paradigm and adapt it for human-centric view synthesis by enabling higher image resolutions, multi-view outputs and stronger cross-view consistency in a single forward pass. We train on a synthetic dataset of human faces spanning diverse identities and apparel. Contrary to diffusion models, this paradigm does not need 2D pre-training and, thanks to its next-scale architecture, it benefits from lower-resolution, general-purpose pre-trainings, with the full-sized purpose-specific images being used only in the last training stages. This enables our architecture to converge with a smaller amount of purpose-specific training data, allowing us to use a smaller but more realistic training dataset. The resulting model produces sharp and realistic views, with the option to synthesize multiple novel viewpoints simultaneously for improved agreement across views. Empirically, we observe gains in perceptual fidelity and cross-view coherence on human subjects, demonstrating that next-scale autoregression is an effective backbone for scalable, multi-output human view synthesis. We also couple our pipeline with an existing transformer-based model for pixel-aligned 3D gaussian lifting from multi-view facial inputs, resulting in accurate and photorealistic 3D models of human faces.