Sphere Encoder 2 improves image generation quality and speed
Sphere Encoder 2
Computer Vision and Pattern Recognition
Summary
Generating images from random points inside a sphere is tricky because points tend to cluster away from the poles, creating gaps in those areas. Also, the original method tries to make images match pixel-for-pixel, leading to blurry results. The authors fixed these problems by adjusting how the model trains and where it samples points, which makes it create clearer images faster.
What this means in practice
- •For computer vision engineers: Create faster and higher-quality generated images for visual tasks using a sphere-based latent representation.
- •For graphic designers: Use improved image generation methods that produce sharper images without complex training overhead.
Authors
Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia, Tom Goldstein
Abstract
Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{https://github.com/kaiyuyue/sphere2}{github.com/kaiyuyue/sphere2}.