Multi subject video editing improves with mask depth and noise control
MDN-Control: Mask-Depth-Noise Guided Region Control for Multi-Subject Video Editing
Computer Vision and Pattern Recognition
Summary
Editing videos with multiple people can be tricky when subjects overlap or cover each other. The paper's authors developed MDN-Control, a method that uses masks to find targets accurately, depth info to handle overlaps, and special noise patterns to start the editing. Their approach helps keep changes only on intended subjects while maintaining video consistency. Tests show their method works better than others on videos with multiple subjects.
What this means in practice
- •For video editors: Edit scenes with several overlapping people accurately while keeping non-target content intact.
- •For visual effects teams: Maintain consistent appearance and boundaries of multiple characters during multi-subject video edits.
Authors
Jiayi Yu, Xi Ye, Lina Wang, Yunkun Xia
Abstract
Multi subject video editing modifies designated subjects while preserving non target content, but faces cross subject attribute leakage, and occlusion ambiguity. Existing approaches rely on masks and struggle to distinguish overlapping subjects or ensure consistent generation. To address these limitations, we propose MDN-Control, a training free framework jointly controlling target localization, occlusion geometry, and appearance initialization. Specifically, mask-guided localization provides consistent target localization, while depth-aware occlusion control resolves ambiguous boundaries between overlapping subjects. We further introduce noise latent prompting, which retrieves Gaussian initializations from a noise library for prompt relevant priors. Experiments on MSVBench show that MDN-Control achieves the lowest CM-Err and the highest Q-Edit, while maintaining competitive text alignment and temporal consistency, demonstrating the effectiveness of combining spatial, geometric, and latent priors for multi subject video editing.