Text-Driven Artistic Staging: Pose, Lighting, and Camera References from Paintings
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
The authors created a new way to generate 3D scenes that include human poses, lighting, and camera angles all at once based on emotion-related text descriptions. They collected data from thousands of paintings by digitally recreating the figures, lighting, and camera views, paired with emotional captions. Then, they trained a model that can take a text description and output different 3D setups matching the described feelings. Their system performed better than previous methods at matching text to scenes while keeping variety. This shows it’s possible to create editable 3D scenes driven by emotional text.
3D staginghuman poseilluminationcamera placementtext-to-3D generationSMPL modelflow-matching transformerArtEmis datasetemotional descriptionretrieval metrics
Authors
Yunge Wen
Abstract
Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model these elements independently. We introduce text-to-editable 3D staging, a task that jointly generates human poses, a dominant light, and a camera configuration from an affective description. We construct 11,911 text--staging pairs from 2,328 figurative paintings by reconstructing SMPL bodies, estimating low-frequency illumination, recovering camera parameters, and pairing each scene with ArtEmis descriptions. We train a flow-matching transformer that supports variable numbers of figures and produces multiple staging alternatives for each prompt. On held-out descriptions, the model achieves 32.2\% retrieval R@1, compared with 16.6\% for CLIP-based nearest-neighbor retrieval, while approximately preserving corpus-level diversity. These results demonstrate the feasibility of generating editable, emotionally conditioned 3D staging references from text.