Papers for

video game animators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Generate realistic listener reactions for natural video conversations

GLARE: Generating Listening Heads with Appropriate Reactions

Abstract: While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.

Wed 30 SeptComputer Vision and Pattern Recognition
The gist
People talking to each other often react with head nods, smiles, or surprised looks, but making computer-generated videos that show these real reactions is hard. The authors created a big new dataset of videos showing listeners’ reactions and labeled exactly when and how they react. They built a system called GLARE that listens to speech sounds and generates matching facial reactions, like nodding or laughing, at the right times. They also made new ways to check if these reactions happen naturally and look right. Their results show it's important to teach computers when and how to react, not just to make faces that look real.
Open → 2609.40317v1

Camera motion direction and speed enhance cinematic shot tools

Unveiling the Value of Motion for Cinematic Camera Trajectories

Abstract: Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGEN, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.

Wed 30 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Camera movement is key to telling stories in movies, not just where the camera is but how it moves. The authors found that describing camera paths using direction and speed works better than just using static camera positions. They created a new way to measure how well this matches descriptions in text and built a tool called CineGEN that makes better camera movements from text. They also made a special movie clip collection with descriptions to test these ideas. This approach helps capture the artistic intent behind camera motion more clearly.
Open → 2609.38683v1