Octree method speeds video processing with fewer data points

Octree-based Video Representation

Computer Vision and Pattern Recognition

Summary

Videos have lots of repeated and simple parts mixed with detailed sections, but current methods treat all parts the same, making processing inefficient. The authors introduce OctVideo, a way to represent video using a tree-like structure that breaks up space and time so simple areas are stored coarsely and detailed parts finely. This method uses fewer computations, encodes and decodes faster, and still produces good quality video reconstructions. It also works well on different video datasets without extra training and helps recognize video content using fewer inputs.

What this means in practice

  • For video processing engineers: Accelerate video encoding and decoding by selectively storing complex video regions more finely, reducing computation and improving speed.
  • For machine learning engineers: Improve video understanding tasks by using fewer video tokens with a more efficient representation for training from scratch.
  • For video streaming service developers: Enable faster video compression and streaming with fewer computational resources, supporting high-quality playback at lower cost.$Commercial implications: This paper’s method allows video streamers to reduce encoding costs and bandwidth usage while maintaining quality, supporting competitive streaming products.

Authors

Rungui Zhou, Chuanzhi Zhou, Yuk-Kit Hou, Peng-Shuai Wang

Abstract

Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.