MatLoom creates compact layered programs for detailed material design
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and LanguageMultimedia
Summary
Creating digital materials for 3D surfaces is tricky because you want not just a picture, but the rules behind how it looks and feels. The authors present MatLoom, a way to turn text descriptions into short, layered programs that define how materials look and interact with light. These programs can be run to produce material maps and can be easily edited later. Without needing special training, MatLoom can improve its designs by checking and fixing errors and tweaking inputs. Tests show MatLoom’s results better match text prompts than leading diffusion methods and users prefer its outputs.
What this means in practice
- •For computer graphics artists: Create editable material assets from text prompts that include appearance rules and layering for detailed control.
- •For game developers: Generate compact material programs from text to efficiently produce and modify realistic surface appearances in games.
Authors
Anson Y. Lam, Shuqing Li, Michael R. Lyu
Abstract
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.