Papers for

video ai engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Improving autoregressive video generation by learning directly from reference videos

From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

Abstract: Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Generating videos one frame at a time often requires complex tools that can be slow and hard to train. The authors developed a way to learn video patterns directly from example videos, skipping some usual extra steps and models. Their method compares videos in a special way that measures differences without needing additional teacher models. This approach improves video quality while keeping the process efficient and allows learning new visual styles and concepts more easily.
Open → 2609.35491v1

Tessera speeds video transformer attention by decoupling mask and GPU tasks

Decoupling Logical Masks from GPU Execution for Dynamic Block-Sparse Attention

Abstract: Attention computation makes inference expensive in video diffusion transformers (vDiTs), which generate videos through iterative denoising. Block-sparse attention (BSA) reduces this cost by computing only blocks selected by a logical mask, which specifies attention interactions to compute. However, coupling logical block geometry to execution choices limits adaptation to varying masks and graphics processing units (GPUs), while runtime kernel specialization can incur preparation overhead that outweighs execution time savings. We present Tessera, a specialized runtime for dynamic BSA that decouples logical masks from GPU execution while preserving specified attention interactions. Its physical mapping layer retains, combines, or subdivides logical attention blocks into physical tiles suited to different attention mask shapes and GPU architectures. Its task organization layer groups and schedules tiles within GPU tasks to reuse data, expose parallelism, and overlap data movement with computation. Finally, profile-guided regime selection enables low- overhead execution plan selection through a lookup table constructed from offline profiling. We implement Tessera with specialized CUDA kernels supporting four NVIDIA GPU generations. Evaluated on 2,315 real attention masks and industrial video diffusion models, Tessera achieves up to 6.79x BSA request speedup over baseline systems in the evaluated video diffusion models.

Tue 22 SeptHardware Architecture
The gist
Video diffusion transformers, which create videos by gradually refining images, spend a lot of time calculating attention, a process that figures out how parts of the video relate to each other. The authors address this by improving block-sparse attention, which skips unnecessary calculations by focusing only on important blocks defined by a logical mask. They introduce Tessera, a system that separates how these blocks are logically defined from how they actually run on GPUs, making it adaptable to different GPU types and attention patterns. Tessera organizes work on the GPU efficiently and picks execution plans quickly, resulting in big speed improvements for video generation models.
Open → 2609.25869v1

Framework measures explanation quality over time in heart ultrasound AI

A Quantitative Evaluation Framework for Temporal Explainability in Echocardiographic Video Segmentation

Abstract: Deep learning has achieved state-of-the-art performance in echocardiographic video segmentation, with an increasing number of models incorporating temporal information. However, quantitative evaluation of temporal explainability remains largely unexplored. We propose a quantitative framework for evaluating Grad-CAM explanations using four complementary metrics measuring temporal consistency, saliency motion, anatomical overlap, and temporal overlap. Using EchoNet-Dynamic, we compare a baseline 2D U-Net with ConvLSTM U-Net models trained across multiple temporal strides. While segmentation performance remained comparable across all models, intermediate ConvLSTM explanations exhibited substantially lower saliency consistency and greater centroid motion than final prediction explanations. Temporal Bottleneck explanations were significantly more stable than Encoder Bottleneck explanations across all strides, while final ConvLSTM Decoder3 explanations were broadly comparable to those of the 2D U-Net. Importantly, conventional frame-wise explanation metrics cannot determine whether variation in intermediate explanations reflects meaningful temporal feature evolution or explanation instability. These findings establish a preliminary quantitative framework for temporal explainability and motivate temporal-aware XAI methods that explicitly account for evolving representations in medical video models.

Mon 7 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Deep learning models can identify heart structures in ultrasound videos, but it’s unclear how well explanations show what the model focuses on over time. The authors created a way to measure how consistent and meaningful these explanations are throughout the video sequence. They found that some model parts produce more stable explanations than others, but common methods can’t always tell if changes in explanations are meaningful or just noise. This work helps guide better explanation tools for medical video AI.
Open → 2609.08043v1