Large dataset enables better editing of versatile 3D digital models
Scaling Versatile 3D Assets Editing with a Million-Scale Dataset
Computer Vision and Pattern Recognition
Summary
Editing 3D digital models is hard because there isn't enough training data or good tools to know where to change the model. The authors created Alchemy3D, a big dataset with over a million 3D models and examples of how to edit them. They trained new AI models that can edit 3D assets based on images or text, learn from few examples, and even help break down models into parts. They also made a new benchmark to test 3D editing tools and showed their approach works better than earlier methods in keeping details, editing accurately, and looking good.
What this means in practice
- •For 3d game developers: Use large-scale trained models to efficiently customize 3D assets based on text or images, improving game content creation.$Commercial implications: Enables commercial game studios to create customizable 3D content faster with better quality using AI-trained editing tools.
- •For animation studios: Assist artists by providing AI tools that edit 3D models with few steps while preserving original details for animation production.
Authors
Badi Li, Tianxin Huang, Yu Zhou, Wei-Shi Zheng, Yi Ma, Shenghua Gao
Abstract
Although recent 3D generative models produce increasingly realistic assets, controllable 3D asset editing remains challenging. Existing methods are limited by scarce training data, insufficient source-aware modeling, and a lack of practical evaluation protocols. To address these limitations, we present Alchemy3D, a unified framework for training and evaluating versatile 3D asset editors that covers data construction, model architecture, and benchmark evaluation. Specifically, we curate Alchemy3D-1M, a large-scale 3D editing dataset containing 1.25M assets and 1.38M editing pairs across seven editing types. On this data, we train a family of generative flow models for general-purpose 3D asset editing. The model family supports image- and text-conditioned editing, few-step inference, and transfer to multi-view 3D part segmentation. We further introduce GEdit3D-Bench, a large-scale, open-world benchmark with a multi-dimensional evaluation protocol. Across existing and newly introduced benchmarks, our method outperforms prior methods on most metrics of editing fidelity, source preservation, and visual quality.