Papers for

video streaming service developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Octree method speeds video processing with fewer data points

Octree-based Video Representation

Abstract: Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Videos have lots of repeated and simple parts mixed with detailed sections, but current methods treat all parts the same, making processing inefficient. The authors introduce OctVideo, a way to represent video using a tree-like structure that breaks up space and time so simple areas are stored coarsely and detailed parts finely. This method uses fewer computations, encodes and decodes faster, and still produces good quality video reconstructions. It also works well on different video datasets without extra training and helps recognize video content using fewer inputs.
Open → 2609.33100v1

Semantic IDs improve item recommendation by tracing interests

From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation

Abstract: Semantic IDs (SIDs) encode each catalog item as a short token sequence, enabling generative recommenders to predict the next item autoregressively. Reasoning-enhanced variants, an increasingly common extension, first generate a textual trace and then decode a next-item SID by beam search. Such recommenders are commonly trained with group-relative policy optimization under an exact-match SID reward, which is sparse in large catalogs. Two failure modes follow. When all rollouts in a group miss the target, the group yields zero advantage and no learning signal. Rollouts sharing the same SID reward receive identical advantages, however much their traces differ. In both cases the reward reflects only the decoded SID, never the reasoning that produced it. This creates a credit-assignment gap. We address this gap with retrieval-grounded query attribution. Each trace is structured into a history summary, a set of interest hypotheses, and a final SID. A frozen retriever executes every hypothesis as a catalog query, so that each hypothesis becomes independently verifiable rather than judged only through the final SID. A rollout is rewarded when any of its queries retrieves the target within the \mbox{top-$K$}, and per-query hit indicators localize that reward to individual hypotheses. Credit is thus assigned at the span level: only hypotheses that individually hit receive positive retrieval advantage, while the retrieval channel never updates the final SID span. Rollouts that share a SID reward can therefore receive different updates. Across experiments on three Amazon Reviews datasets, this yields consistent improvements in SID recommendation. On Video Games, an oracle analysis further reveals the potential of interest-conditioned SID decoding: selecting the target-relevant query among generated interests improves both recall and ranking.

Thu 24 SeptInformation RetrievalArtificial Intelligence
The gist
The paper focuses on improving how recommendation systems predict the next item a user might want to see, like in shopping or streaming services. Traditional methods only reward correct item predictions but ignore the reasoning process that led to those predictions, which limits learning. The authors introduce a method that breaks down the reasoning into smaller parts and checks which parts actually helped find the right item, giving more precise feedback. This approach helps the system learn better to recommend items by connecting interests to actual results.
Open → 2609.29983v1

New method improves long sequence item recommendations with better speed and accuracy

Preference-Drift-Aware Subsequence Learning and Hierarchical Context Fusion for Long-Sequence Generative Recommendation

Abstract: Long-sequence generative recommendation methods autoregressively model the user's interaction sequence to generate the next-item representation. Existing methods generally fall into two categories: efficient full-sequence modeling and target-aware context retrieval. Our experiments reveal that as the sequence length increases, the former incurs steadily growing computational cost while its accuracy gains quickly saturate and even degrade due to noise; the latter, though shortening the input sequence, is susceptible to noise that is semantically consistent yet preference-inconsistent, as well as to incomplete contexts. Both paradigms ignore the dynamic changes of user preferences and the cross-subsequence dependencies when handling historical information, thereby limiting accuracy and efficiency. To address these issues, we propose a preference-drift-aware subsequence learning and hierarchical context fusion for long-sequence generative recommendation. Specifically, we learn differentiable soft subsequence boundaries using multidimensional preference-drift information and aggregate items within each subsequence into preference-coherent representations via linear attention with soft assignment weights, thereby circumventing the expense of full-sequence attention. A cross-attention mechanism is then employed to capture dependencies between recent interactions and relevant subsequence contexts, mitigating noise in learning recent-item representations. Finally, a gated fusion mechanism adaptively combines the recent-item representation with the global subsequence context, allowing the resulting target representation to encode both recent and long-term preferences. Extensive experiments demonstrate that our method consistently outperforms existing baselines in both recommendation accuracy and computational efficiency.

Fri 11 SeptInformation Retrieval
The gist
Many recommendation systems try to guess what you’ll like next by looking at everything you've done before, but this can be slow and sometimes less accurate when there’s a lot of history. The authors found that current methods either take too long with long histories or don’t handle changes in your tastes well. They created a new approach that splits your activity into smaller chunks and combines recent and long-term preferences in a smart way. This makes recommendations both faster to compute and better at matching what you want now.
Open → 2609.12556v1