Papers for

visual effects teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

MaLiang harness bridges code and visuals for better image and video generation

MaLiang-Harness: A Programmable Path to Image and Video Generation

Abstract: Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Creating images and videos from computer programs can be tricky because the final picture might not match the instructions exactly. The authors of this paper identified this gap between what a visual program should do and what it actually produces, calling it the Program-to-Visual (P2V) gap. They built MaLiang-Harness, a system that helps create visuals by letting machines construct, check, and improve code step-by-step, keeping track of changes and verifying results as they go. They tested this system with multiple powerful language models and found that their approach helps ensure visuals match what the code intends, revealing important insights about how these models work with visual tasks.
Open → 2609.34309v1

Memory methods improve autoregressive video generation over time

The Past Frames the Future: Memory for Autoregressive Video Generation

Abstract: Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, practical models must operate under strictly bounded context windows, storage, and computational limits. Consequently, critical historical information, e.g., entity identities, dynamic states, and intervention-induced causal changes, often leaves the active context long before its relevance diminishes. Overcoming this limitation and maintaining temporal persistence constitutes a fundamental memory problem. We present a systematic and comprehensive review of memory mechanisms in AR video generation. We formulate memory operationally as persistent historical information maintained across outer AR steps, capable of influencing future generation even after the originating evidence is no longer locally accessible. Building upon this unified framework, we organize the literature through five complementary perspectives: (I) Forms, the representational carriers of history; (II) Functions, the specific semantic and physical information requiring preservation; (III) Operations, the lifecycle of writing, reading, updating, managing, and integrating memory; (IV) Learning, the optimization of memory behaviors under closed-loop rollouts; and (V) Evaluation, the paradigms for diagnosing genuine memory capabilities. We conclude by synthesizing open challenges, including composable and resource-aware memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation. By bridging representations, mechanisms, and learning paradigms, this paper establishes a structured foundation for developing reliable, memory-conditioned video generation systems.

Wed 23 SeptComputer Vision and Pattern Recognition
The gist
Generating videos step-by-step faces a problem when past details get lost after a short time, making it hard to keep things consistent like the identity of objects or changes caused by actions. This paper reviews different strategies for keeping memory during video generation, explaining how information is stored, updated, and used to improve videos over long periods. The authors also discuss how to teach these systems to remember better and how to test if memory is really being used. Their work helps create systems that generate longer, more coherent videos by remembering important past details.
Open → 2609.28466v1

Text to video system generates physically realistic complex motions

CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation

Abstract: While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a physics-consistent T2V framework for adaptable generation across complex real-world scenarios. It extends neural dynamics modeling beyond single-type motions to encompass composite physical behaviors, including coupled motions, multi-stage transitions, and multi-object collisions. Furthermore, CompAdapt translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial physical parameters. To generalize to novel physical environments, CompAdapt introduces dynamics-aware prior matching, achieving one-shot adaptation without retraining the core dynamics module. In addition, a physics-aware latent feature fusion module improves visual fidelity under fast and complex motion. Experiments on physics-focused T2V benchmarks demonstrate that CompAdapt improves physical consistency over both general T2V models and physics-constrained baselines, while preserving high visual quality and adaptability to unseen dynamics. The project page is available at https://makapic.github.io/CompAdapt/ .

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
Videos created from text prompts often look smooth but don’t follow real-world physics. The authors designed a system called CompAdapt that can create videos respecting complex physical motions like collisions and multi-stage actions. It turns language instructions into detailed motion rules and adapts quickly to new kinds of physical behavior without re-training. Their experiments show this method makes videos that look better and act more physically correct than earlier systems.
Open → 2609.21455v1

Iterative refinement improves dynamic 4D scene synthesis from sparse videos

4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors

Abstract: This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
Creating detailed 3D videos from just a few camera views is difficult because it’s hard to guess parts of the scene that aren’t seen. The authors developed a method that first builds a fuller 3D shape from sparse views, then uses AI to improve the video frames and fix mistakes step by step. This approach helps fill in missing details and makes the dynamic 3D scenes much clearer and more complete. Their method showed significantly better results than earlier techniques on a standard test.
Open → 2609.21176v1

Multi subject video editing improves with mask depth and noise control

MDN-Control: Mask-Depth-Noise Guided Region Control for Multi-Subject Video Editing

Abstract: Multi subject video editing modifies designated subjects while preserving non target content, but faces cross subject attribute leakage, and occlusion ambiguity. Existing approaches rely on masks and struggle to distinguish overlapping subjects or ensure consistent generation. To address these limitations, we propose MDN-Control, a training free framework jointly controlling target localization, occlusion geometry, and appearance initialization. Specifically, mask-guided localization provides consistent target localization, while depth-aware occlusion control resolves ambiguous boundaries between overlapping subjects. We further introduce noise latent prompting, which retrieves Gaussian initializations from a noise library for prompt relevant priors. Experiments on MSVBench show that MDN-Control achieves the lowest CM-Err and the highest Q-Edit, while maintaining competitive text alignment and temporal consistency, demonstrating the effectiveness of combining spatial, geometric, and latent priors for multi subject video editing.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Editing videos with multiple people can be tricky when subjects overlap or cover each other. The paper's authors developed MDN-Control, a method that uses masks to find targets accurately, depth info to handle overlaps, and special noise patterns to start the editing. Their approach helps keep changes only on intended subjects while maintaining video consistency. Tests show their method works better than others on videos with multiple subjects.
Open → 2609.16475v1

PhysFlow generates videos with more realistic motion and appearance

PhysFlow: Physics-Aware Optical Flow for Motion Controllable Video Generation

Abstract: Video generation models have recently attracted substantial attention for their ability to generate visually compelling videos, yet ensuring physically consistent and plausible dynamics still remains a fundamental challenge, driving a growing line of research on physical realism in video generation. To address this challenge, motivated by the fact that physical regularities are primarily encoded in motion patterns, we propose PhysFlow, a novel two-stage framework for improving the physical plausibility of generated videos by decomposing video generation into motion-aware optical flow generation followed by motion-conditioned appearance synthesis. Specifically, PhysFlow consists of a physics-aware optical-flow video generator called PA-Flow and a flow-guided video generator called FlowRender. During the first stage, PA-Flow employs a physics-aware attention module to model how motion attributes and material properties influence global motion and local deformation, respectively, and generates an optical flow video as an explicit representation of motion. In the second stage, FlowRender leverages the decoupled motion representation as guidance to synthesize realistic textures and appearances, ultimately producing the final physically plausible video. To further support model training with explicit physical supervision, we construct PhysVideo, a physics-based video dataset generated with a physics engine and 3D-GS rendering, containing 10K foreground objects and 50K realistic video sequences with annotations of motion and material properties. Extensive experiments demonstrate that our proposed PhysFlow generates videos with superior physical plausibility while maintaining high visual fidelity compared with existing methods.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Videos created by computers often look nice but can have weird or impossible movements. The researchers designed PhysFlow, which first predicts how things in a video move with physics rules, then uses those movements to create detailed video images. They also made a new dataset of videos created with physics engines to help train and test the system. PhysFlow produces videos that look both good and physically believable compared to other methods.
Open → 2609.08215v1