Text to image generation lets users control fine grained emotions

Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance

Computer Vision and Pattern Recognition

Summary

Text-to-image tools usually let you describe what you want to see but not the exact feelings an image should show. The authors created a method called EMOTRANS that uses a psychological model of emotions (valence, arousal, dominance) to guide image generation separately from the description. They made a new dataset and trained a model to change emotional style in pictures smoothly, without messing up the content or quality. Their system helps creators adjust the emotional tone of images with more precision and clearer intensity levels.

What this means in practice

  • For graphic designers: Create images where the emotional tone can be finely tuned separately from the content description, enabling nuanced artistic expression.
  • For advertising teams: Produce visual ads that reflect precise emotional states to better target viewer feelings and intentions without changing core messages.

Authors

Minglang Li, Yueyue Fang, Xieping Gao

Abstract

Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should convey without rewriting its content description. Natural language can suggest emotions, but it offers no control scale with stable meanings and ordered intensities. We propose EMOTRANS, which transforms psychologically grounded valence-arousal-dominance (VAD) coordinates into generation conditions that are independent of the content text and modulated across denoising stages, making emotional style a finely adjustable creative variable. To support this goal, we construct EMOVAD, an art-painting dataset that pairs objective content descriptions with separately collected emotional ratings from multiple annotators. We also coordinate emotional expression and content preservation through dual-branch training with a shared model. Objective and human evaluations show that the framework improves the accuracy of three-dimensional emotion control and produces perceptible, orderable continuous changes while maintaining competitive text alignment and image quality. This work provides a practical emotion-driven approach to image generation that extends objective content depiction to fine-grained emotional adjustment.