Trigger indexed memory improves personalized text to image generation
V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization
Artificial Intelligence
Summary
Pretrained text-to-image models are good at generating images from descriptions but struggle to accurately create specific characters or objects from just a few examples. The authors present V-Engram, a method that stores detailed visual info separately and links it to special keywords or 'triggers.' This helps the model remember and use specific appearances without changing its main knowledge, allowing for multiple personalized items in one prompt and better control. Their tests show V-Engram matches other methods in image accuracy while being more flexible and efficient.
What this means in practice
- •For digital artists: Create personalized characters or objects in images from few references without retraining the entire model.
- •For content creators: Manage multiple customized visual concepts in a single text prompt for flexible image generation.
Authors
Haoran He, Runyuan Cai, Yiming Wang, Lin Yu, Xiaodong Zeng
Abstract
Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model's base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.