Papers for

video editors

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Reusable personalized color editing method adapts photo look quickly

PrefLUT: Reusable and Refinable Personalized Color Editing from Pairwise Preferences

Abstract: Photographic color editing is inherently personal: the same image can appear too warm, too muted, or already satisfactory to different users. Most lookup table (LUT) and reference-guided methods target a specified appearance rather than model persistent preferences from repeated user choices. To address this gap, we introduce PrefLUT, a reusable and refinable user-preference modeling framework for deployable 3D LUTs, encoding ordered preferred/non-preferred image pairs into a lightweight Reusable User Profile that is reused across queries and refined using additional user preference pairs, without per-user optimization. A Query-Conditioned LUT Predictor combines this profile with each image to predict a LUT latent vector and edit strength. An Identity-Residual LUT Decoder and Edit-Strength Controller then produce an exportable 3D LUT. Experiments on three datasets demonstrate effective personalized editing and general-purpose enhancement. Each quantized profile requires only 260 bytes, and editing takes 1.365 ms/image on an RTX 5090 GPU. We also introduce the Preference-Conditioning Verification Protocol (PCVP), an evaluation protocol to verify whether personalized image edits depend on user preferences and the query image through controlled changes to user profiles, preference orders, pair correspondences, and query images.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
People often like different color edits on the same photo, but usual editing tools target just one fixed look. The authors created PrefLUT, which learns a user’s color preferences from pairs of images they like or dislike, stores these in a tiny profile, and reuses or updates it easily to edit new photos. This approach works fast and efficiently on powerful GPUs and can create standard color lookup tables used in photo editing software. The authors also designed a test to make sure the edits truly reflect user preferences and the pictures involved.
Open → 2609.34133v1

CraftTrace enables flexible editing of full videos using video structures

CraftTrace: Unflattening Videos into Malleable, Creation-Inspired Structures for Generative Editing

Abstract: Recent generative video editing models enable video content modification (e.g., changing a character) but target short clips. Extending them to full multi-shot videos requires tedious work to locate relevant content across shots, segment it into clips, craft context-aware editing prompts for each clip, and repeatedly articulate complex editing intent. To address this, we explore an interaction paradigm for editing through underlying video structures (e.g., scripts, scenes, characters, shots, and their relationships). We present CraftTrace, an interactive prototype that transforms a video into a malleable, multilevel structure for generative editing. Users work in task-centric workspaces to modify elements or reshape relationships, while an AI agent translates and propagates changes across the video. A user study and expert review show that this structure helps users understand videos, formulate and refine editing intent, and explore alternatives, supporting rapid prototyping during early-stage exploration and full video post-production.

Thu 24 SeptHuman-Computer Interaction
The gist
Editing videos with many scenes and shots is hard because you must find and change parts one by one. The authors introduced CraftTrace, a tool that turns a whole video into a map of scenes, characters, and shots you can easily change. An AI helps apply your edits across the whole video so you don’t have to do repetitive work. This makes exploring ideas and refining edits faster and easier.
Open → 2609.30623v1

SSE model enhances and remixes audio using video and text guidance

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

Abstract: We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site

Thu 24 SeptSoundArtificial Intelligence
The gist
Mixing audio in videos is hard when unwanted sounds or echo get in the way. The authors created a new system called SSE that uses the video and text descriptions to find, separate, and improve sounds in a video’s audio track. This lets people remove unwanted noise, balance sound levels, and reduce echoes with simple instructions. They also made a new dataset to help train and test their system and showed it works better than existing methods.
Open → 2609.29169v1

Lightweight model separates singing voices using audio and video cues

MambaVoice: Lightweight Audiovisual Singing Voice Separation Via A Hybrid Mamba-Transformer Model

Abstract: Isolating a target singing voice from a music video remains challenging, particularly in the presence of multiple vocalists and dense instrumental accompaniment. We propose MambaVoice, a lightweight audiovisual framework that leverages a hybrid Mamba--Transformer architecture for targeted singing voice separation. The model jointly encodes audio and visual streams using an attention-based band-split audio encoder and a spatio-temporal graph convolutional network (ST-GCN) for facial motion features. These modalities are fused through a multiplicative gating mechanism, enabling visual cues to selectively modulate audio representations. The fused features are processed by a hybrid backbone that combines Transformer self-attention with Selective State Space Models (SSMs), achieving efficient long-range temporal modeling with linear complexity. We evaluated MambaVoice on the Acappella and URSing datasets under challenging conditions, including mixtures with interfering singers. At 16.2 million parameters, the model demonstrates comparable performance, achieving 14.18 dB SDR on Acappella and strong cross-dataset performance on URSing, comparable to larger models at a fraction of the parameter count. These findings highlight the effectiveness of hybrid SSM--attention architectures for scalable, efficient audiovisual source separation, suggesting they are well-suited as lightweight components within larger pipelines. We conduct a perceptual study that further supports our improvements in objective metrics. We provide our implementation online.

Tue 22 SeptSound
The gist
Separating a singer’s voice from music videos is hard, especially when many people sing or lots of instruments play. The authors created MambaVoice, a compact model that uses both sound and face movements to pick out the target singer’s voice. It combines audio and video information in a clever way to focus on the right voice, and uses advanced methods to understand long stretches of music efficiently. Tests show MambaVoice works well compared to bigger models but with fewer resources.
Open → 2609.26635v1

Visual and text guided method removes sounds linked to video objects

TV-AudioRemover: Joint Text-Visual Guided Sound Removal with Multi-Task Hard-Mixture Curriculum

Abstract: Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and therefore rely on limited single-modal control, which is less effective than multimodal guidance that provides stronger semantic grounding and temporal synchronization cues. In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture. To acquire high-quality training data, we devise a pipeline to construct a million-scale dataset of single-object audio-visual aligned samples, from which we synthesize mixture-target pairs customized for model training. To effectively leverage visual context and follow instruction intent, we augment the model architecture with task tokens, generalizable instruction modeling, and modality-specific global guidance. We further adopt multi-task training to strengthen task-role comprehension, and employ a hard-mixture curriculum that leverages semantically similar acoustic mixtures during fine-tuning to enhance fine-grained source discrimination. To support evaluation, we present AV-Remove-Bench, a comprehensive audio-visual object removal benchmark, along with dedicated objective metrics and an MLLM-based evaluation protocol. Experiments demonstrate that our method achieves state-of-the-art performance on both subjective and objective metrics. Project page: https://yjx-research.github.io/TV-AudioRemover/.

Tue 22 SeptMultimediaComputer Vision and Pattern RecognitionSound
The gist
When people remove objects from a video, the sounds those objects made often remain, making the video feel strange. The authors created a method that uses both the edited video and text instructions to remove the sounds linked to the removed objects. They trained their model on a large set of video and audio clips with aligned objects and sounds. Their approach improves how well the unwanted sounds are canceled while keeping the rest intact. They also created a new benchmark to measure how well these methods work.
Open → 2609.25864v1

Qwen3.8-Omni improves multimodal AI for long tasks and video work

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Abstract: We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash substantially improves multimodal understanding and reasoning, as well as performance on long-horizon agentic tasks. These capabilities are supported by a native multimodal co-training strategy that preserves strong text-domain capabilities while facilitating the transfer of agentic capabilities from text to audio and video tasks. The model inherits the sparse mixture-of-experts (MoE) architecture of Qwen3.8-Next and extends the context window to one million tokens, supporting long-context multimodal reasoning and long-horizon planning. These advances enable integration into production workflows as a primary agent or a specialized sub-agent, supporting video editing, long-form audio and video translation, music-conditioned music video or movie generation, and video-based note or omni-skill creation. To address the lack of native audio and video support in existing agent harnesses, we release Qwen-MM-Plugins, a lightweight open-source plugin framework for multimodal productivity. We further frame real-time multimodal interaction as a system-level challenge requiring orchestration of context and memory management, tool use, and sub-agent delegation. Accordingly, we release Qwen-Live-Harness, an open-source framework for building responsive, real-time multimodal agents based on Qwen3.8-Omni-Flash. Extensive evaluations demonstrate that Qwen3.8-Omni-Flash achieves strong performance across multimodal understanding, reasoning, long-horizon agentic execution, and video productivity tasks. These results and the accompanying open-source tools support Qwen3.8-Omni-Flash as a practical foundation for deploying natively multimodal agents in research and production.

Tue 22 SeptComputation and LanguageComputer Vision and Pattern RecognitionMultimedia
The gist
Handling different types of information like text, audio, and video at once is hard for AI. The authors developed Qwen3.8-Omni-Flash, an AI model that can understand and reason about multimodal data over very long conversations or projects. It also helps with tasks like video editing and music video creation. They created new software tools to support real-time interaction and seamless work across different media types. This model and tools aim to make AI assistants better at complex, long-lasting tasks involving various media.
Open → 2609.25611v1

Edit VAR improves text guided video editing with better source and motion control

Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

Abstract: Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
Editing videos by changing their content based on text is tricky because you need to keep parts of the video unchanged and consistent over time. The authors introduce Edit-VAR, a method that does this without needing extra training or reversing the video editing process. It smartly replaces parts of video data based on text, keeping unedited parts intact and making sure moving details stay smooth. Tests show it works better than existing methods in making edits that look good, keep the original video’s feel, and run efficiently.
Open → 2609.21268v1

Multi subject video editing improves with mask depth and noise control

MDN-Control: Mask-Depth-Noise Guided Region Control for Multi-Subject Video Editing

Abstract: Multi subject video editing modifies designated subjects while preserving non target content, but faces cross subject attribute leakage, and occlusion ambiguity. Existing approaches rely on masks and struggle to distinguish overlapping subjects or ensure consistent generation. To address these limitations, we propose MDN-Control, a training free framework jointly controlling target localization, occlusion geometry, and appearance initialization. Specifically, mask-guided localization provides consistent target localization, while depth-aware occlusion control resolves ambiguous boundaries between overlapping subjects. We further introduce noise latent prompting, which retrieves Gaussian initializations from a noise library for prompt relevant priors. Experiments on MSVBench show that MDN-Control achieves the lowest CM-Err and the highest Q-Edit, while maintaining competitive text alignment and temporal consistency, demonstrating the effectiveness of combining spatial, geometric, and latent priors for multi subject video editing.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Editing videos with multiple people can be tricky when subjects overlap or cover each other. The paper's authors developed MDN-Control, a method that uses masks to find targets accurately, depth info to handle overlaps, and special noise patterns to start the editing. Their approach helps keep changes only on intended subjects while maintaining video consistency. Tests show their method works better than others on videos with multiple subjects.
Open → 2609.16475v1