SkillPE evolves cinematic skills for better text-to-video prompts

SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Creating great cinematic videos from text prompts is hard for people who are not experts in filmmaking. The authors developed SkillPE, a framework that improves prompts by using filmmaking skills like shot composition and lighting, learned from expert examples. SkillPE also looks at movie references that match or creatively differ to help generate better video results. This approach helps balance how well the video matches the prompt and how creative or cinematic it looks. Tests show SkillPE improves video quality and storytelling compared to other methods.

What this means in practice

  • For video game developers: Enhance in-game cinematic scenes by generating creative, high-fidelity video sequences from text prompts using evolved filmmaking skills.
  • For advertising teams: Produce visually compelling promotional videos with improved creative direction from text-to-video tools empowered by cinematic skill evolution.

Authors

Yanwei Huang, Mingxuan Zhu, Shujie Li, Shiyuan Liu, Yuanxing Zhang, Arpit Narechania

Abstract

Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents shot logic, composition, lighting, sound design, and other filmmaking cues in a fine-grained format, and retrieves movie references categorized as resonators (good matches), dissonants (weak matches), and divergents (creatively useful near-misses). The first two refine when and how a skill should be applied, while divergents inspire alternative cinematic realizations at different degrees of modification while preserving the user intent. Candidate skills are assessed through generated videos along prompt fidelity, cinematic quality, narrative appeal, and creativity to construct the final skill libraries. Experiments on StoryEval and VBench show improvements of up to 1.40 points over the strongest external baseline and 0.51 points over seed skills on 7-point four-dimensional evaluation, while remaining competitive on benchmark-native metrics. Overall, SkillPE offers a practical approach to balancing fidelity and creativity in cinematic text-to-video generation. Code is available at https://github.com/Ais0n/SkillPE .