Video generation speed improved up to 30 times on cloud and edge devices
Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
Artificial IntelligenceComputer Vision and Pattern RecognitionMachine Learning
Summary
Generating videos with advanced AI models is often very slow because they use huge amounts of computing power and memory. The authors developed a technique that speeds up this video creation by changing how the process works: first creating a low-detail version quickly, then refining it in more detail. They also made computer code run faster by automatically optimizing its steps and memory use. These improvements let big videos with sound be made much faster on powerful cloud servers and fit into memory on smaller edge devices.
What this means in practice
- •For cloud infrastructure teams: Deploy faster video generation models by reducing latency and memory use on large multi-GPU nodes.
- •For edge device developers: Run advanced video diffusion models with saved memory enabling real-time generation on limited hardware like DGX Spark.
Authors
Yitong Li, Jincheng Yu, Junsong Chen, Haopeng Li, Shuchen Xue, Haozhe Liu, Ping Luo, Song Han, Enze Xie
Abstract
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.