Interactive humanoid videos created with real time multimodal control

FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning

Computer Vision and Pattern Recognition

Summary

Creating realistic animated videos of people that interact in real time is a challenge because it requires understanding and reacting to many inputs like speech, images, and actions. The authors introduce FlowAct-R2, a system that uses a special technique to take in ongoing audio and images while planning ahead how the animated person should behave. This helps the video avatar maintain a consistent look and respond smoothly to what viewers say or do live. The system can create high-quality videos that run continuously for a long time and fit scenarios like live streaming, shopping, or chatting.

What this means in practice

  • For entertainment streamers: Generate live interactive humanoid videos that respond to audience inputs and maintain consistent appearance in real time.$Commercial implications: Enables live-streaming platforms to create dynamic avatar hosts that enhance viewer engagement and extend session durations.
  • For live shopping hosts: Produce real-time animated hosts that can respond to customer interactions and manage scheduled product showcases effectively.

Authors

Ziyao Huang, Zhengkun Rong, Shiyang Qin, Shuang Liang, Wentao Hu, Yuxuan Luo, Yuan Zhang, Mingyuan Gao

Abstract

We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini reference-to-video backbone to accept rolling action prompts, streaming audio, and dynamically updated image, audio, and video references. Video-driven rotary positional embeddings align reference chunks with the generation timeline, while reference-plus-image conditioning and partially noised historical motion frames preserve appearance and avoid accumulated drift. Second, a Proactive Interaction Agent separates pre-online planning from online scheduling and response: it prepares a persona, a long-horizon agenda, and reusable multimodal skills in advance, then autonomously schedules behaviors, responds to audience input, and handles interruptions during a live session. FlowAct-R2 supports real-time 720p generation and hour-scale streaming across entertainment streaming, live shopping, video chatting, and live vlogging.