Papers for

video software developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Video models contain correct motion data even when output is wrong

A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models

Abstract: When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in this conflicting case, a low-dimensional edit predicted from simple physical variables restores the correct fast motion. We call this ability causal writability. At fixed strength, we find a sharp depth boundary: the same edit changes the video before the boundary but not after it. This closure marks commitment for that write. The motion signal nevertheless remains, and a stronger downstream write can restore physical motion, while excessive gain overshoots. Early causal writability predicts which errors training later corrects: those errors are writable at more network depths than errors that persist. We reproduce both causal writability and its sharp closure in a pretrained 1.3B video model, supporting generality across model scale and training regime.

Mon 14 SeptMachine LearningComputer Vision and Pattern Recognition
The gist
Sometimes video models produce motions that look wrong, and it's unclear if they never learned the right motion or just aren't using it properly. The authors show that the models actually do have the correct motion inside them and that simple edits can make the videos move correctly. They discovered a sharp point in the model where changes stop affecting the output, marking when the model commits to a motion. This finding helps explain which errors can be fixed by further training and applies broadly to large video models.
Open 2609.15980v1

Video-MOPD improves video understanding with multi-teacher learning

Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

Abstract: Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize within a single model. We introduce Video-MOPD-8B, an open-weight model dedicated to video understanding tasks. To fundamentally enhance its capabilities, we conduct targeted reinforcement learning (RL) optimization across three core domains: video temporal grounding (VTG), general video comprehension, and video STEM reasoning. We then unify their complementary capabilities via Multi-Teacher On-Policy Distillation (MOPD), which consolidates expert knowledge by supervising student-generated trajectories with routed teacher feedback. We further introduce Reliability-Aware Informative Sampling (RAIS), which selects examples with consistently reliable teacher supervision and large teacher-student performance gaps. Together, these components enable Video-MOPD-8B to achieve coordinated and comprehensive performance gains across diverse video understanding tasks. Extensive experiments on comprehensive benchmarks covering general video understanding, temporal grounding, video reasoning, and video STEM tasks demonstrate that Video-MOPD-8B achieves state-of-the-art performance among existing models at a comparable scale. The trained model weights are available at https://huggingface.co/LandH/Video-MOPD-8B.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Understanding videos well requires a model to do several tricky tasks at once, like noticing details, understanding timing, and solving problems. The authors created Video-MOPD-8B, a model trained to get better at these by learning from multiple expert models each focused on a different skill. They use a special training method that combines feedback from these expert teachers while picking the best learning examples. This approach helps the model perform well across many types of video tasks.
Open 2609.09300v1