Sparse attention method cuts cost in 3D vision transformers
ReSS: Residual-Restoring Sparse Attention for 3D Vision Transformers
Computer Vision and Pattern Recognition
Summary
3D vision transformers analyze multiple images to understand 3D scenes, but they slow down as the number of images increases because they look at everything at once. The authors found that past methods that remove some parts (blocks) based on attention probability harm performance because this probability doesn’t reflect how much impact those parts have. They created a new way called ReSS that chooses parts to keep by measuring how much each part changes the model's important internal signals, keeping the output closer to the original. Tests show ReSS keeps accuracy better than older methods when reducing work.
What this means in practice
- •For computer vision engineers: Reduce computation in multi-view 3D reconstruction models by sparsifying attention while maintaining accuracy.
- •For augmented reality developers: Build efficient 3D scene understanding systems that run faster on limited hardware by applying selective attention based on residual drift.
Authors
Yongsung Kim, Jaehoon Lee, Minjun Park, Wooseok Song, Hun Hwangbo, Sungroh Yoon
Abstract
3D vision transformers such as VGGT predict camera poses and scene geometry from multi-view images in a single forward pass, but their global attention over all concatenated view tokens dominates computation as the number of views grows. To reduce this cost, SparseVGGT and HeSS sparsify attention at the block level, and both retain blocks with high attention probability. However, we observe that attention probability poorly predicts how much the model's behavior actually changes when a block is removed, and we show that this mismatch is why performance collapses as sparsity increases. In this paper, we propose ReSS (ReSidual-ReStoring Sparse Attention), which recasts block selection from a problem of maximizing the retained attention mass to one of minimizing the drift that sparsification leaves in the residual stream. We introduce a drift score that quantifies how much each block shifts the residual, and, since the drift of a drop set depends on the directions of the contribution vectors rather than on their magnitudes alone, an iterative residual restoration procedure that refines the drop set as a whole. Across three backbones and five datasets, ReSS preserves dense performance better than prior methods at matched sparsity. Two further results support drift as the quantity that governs the cost of sparsification: maximizing drift degrades performance faster than random selection, and plotted against realized drift instead of sparsity, all methods fall approximately onto a single curve. Code is available at https://github.com/libary753/ReSS.