Papers for

media content platforms

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Ai agents struggle with real video editing in finished deliverables

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

Abstract: AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceMultimedia
The gist
Video editing involves many complex steps, from choosing the right shots to assembling a polished final video. The authors created Timeline-Bench, a test set of 56 real video editing tasks, to see how well AI agents can complete these tasks from start to finish. They found that even the best AI agents only completed about one quarter of the tasks successfully, mostly failing the quality checks. Human editors preferred the original edits most of the time, showing that current AI still lacks creativity and finesse in video editing.
Open → 2609.35143v1

Large AI models struggle to understand cinematic storytelling techniques

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

Abstract: Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.

Wed 23 SeptComputer Vision and Pattern Recognition
The gist
Understanding how movies tell stories visually is hard for AI models that see and read videos. The authors created a new test called CinematicVQA, which checks if AI can reason about the film-making methods behind what we see, not just describe the images. They found that AI models are better at describing scenes visually than explaining the storytelling tricks used, and common techniques to improve reasoning did not help. Training AI on this new test helped them get better at understanding narrative and multi-step story reasoning.
Open → 2609.28813v1

VideoX-Qwen enables instruction-based video editing with large data and models

VideoX-Qwen: Data-Centric Instruction-Based Video Editing

Abstract: Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion, and temporal continuity. We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing. Our scalable production pipeline organizes specialized generation and understanding models into complementary routes for addition, removal, replacement, and attribute editing, followed by quality screening and instruction enrichment. It produces more than 1.2 million directional video-editing records, including over 400,000 records in each major task group, with an automatic acceptance rate of 89%. The resulting corpus provides broad and structured coverage of common editing operations through a unified source-instruction-target interface. We further develop a unified Qwen-Wan editor that combines multimodal semantic conditioning with dense source-video latent guidance. A progressive image-video training strategy aligns the multimodal instruction interface, adapts the video generator to source-conditioned editing, and refines output quality with selected high-resolution data. In a 100-example comparison with UniVideo and Kling O1, VideoX-Qwen achieves the best mean result on nine of eleven reported metrics, including instruction following, editing quality, content preservation, structural and perceptual similarity, and video-distribution quality. Together, the large-scale data-production system and unified training framework provide a practical foundation for more capable instruction-driven video editing.

Tue 22 SeptArtificial IntelligenceGraphics
The gist
Editing videos so they follow detailed instructions while keeping the original motion and scenes intact is much harder than just creating videos from scratch. The authors developed VideoX-Qwen, a system that builds an enormous dataset of video editing examples by combining multiple specialized models to add, remove, replace, or change things based on instructions. They then trained a video editing model that understands instructions and maintains video quality. When tested, VideoX-Qwen showed better results than previous methods on many measures like following instructions and preserving content.
Open → 2609.26015v1