SubjectAnchor improves identity consistency in multi-shot video storytelling

SubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot Storytelling

Computer Vision and Pattern Recognition

Summary

Keeping characters looking the same in different parts of a video story is hard when scenes change quickly. The authors present SubjectAnchor, a method that remembers key images of each character from earlier scenes to help the computer generate new scenes that keep characters consistent. It uses a special way to organize and focus on these remembered images so the story looks smoother across different shots. Tests show this method does better than previous ones at keeping characters looking right without losing video quality.

What this means in practice

  • For video content creators: Generate multi-scene video stories with consistent character appearances across shots using shot-wise prompts and visual memory retrieval.
  • For digital animation studios: Maintain visual continuity of characters across multiple scenes in animated video sequences by conditioning generation on visual memories of prior shots.

Authors

Xinyu Wang, Huafeng Shi, Zian Li, Yan Zhou, Xiaoqiang Liu, Yue Ma, Pengfei Wan

Abstract

We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retaining the controllability of shot-wise prompting. Built on Wan2.2-I2V-A14B, SubjectAnchor contains three key components: subject-related memory construction, subject-aware temporal rotary position encoding, and memory-aware attention partition. For each target shot, the method constructs a compact memory bank by tracing each required subject to its historical appearance and retrieving the most relevant precomputed keyframes. These memory frames are encoded into the model input as explicit visual conditions, while different subjects are assigned to separated negative temporal slots to reduce identity interference. In addition, memory-aware attention partition regulates the interaction between memory tokens and generated content within a shared backbone. This formulation preserves the appearance anchoring of explicit visual memory while remaining compatible with script-driven shot-by-shot generation. Experiments show that SubjectAnchor improves cross-shot identity consistency over representative memory-based and holistic baselines while maintaining competitive visual quality.