DORA cuts vision transformer computation while keeping accuracy high
DORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision Transformers
Computer Vision and Pattern Recognition
Summary
Vision Transformers analyze images by looking at many small parts called tokens, but this can be very slow since it requires a lot of calculations. The authors created DORA, a smart agent that learns to remove unnecessary tokens during processing, making the analysis faster without hurting accuracy much. DORA decides dynamically for each image and at multiple steps which tokens to keep or drop, adapting better than fixed token reduction methods. This approach greatly reduces the work Vision Transformers need while keeping nearly the same quality of image recognition.
What this means in practice
- •For computer vision engineers: Implement dynamic token pruning to speed up image classification models while preserving accuracy under varied input complexity.
- •For mobile device developers: Deploy frozen vision transformer models with real-time token pruning to achieve faster inference and lower power consumption on edge devices.
Authors
Kaixuan He, Song Chen, Yi Kang
Abstract
Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose DORA (Dynamic Online Reinforcement Agent), which learns an input-adaptive pruning policy itself for frozen ViTs. At each eligible block, a hierarchical actor decides whether to prune, how many tokens to remove, and which tokens to remove from each image's evolving representation. Because early deletions change the states observed by later decisions, DORA formulates pruning as a finite-horizon Markov decision process. Complete-prefix shadow evaluations convert final-prediction fidelity into localized per-step credit, while closed-loop accuracy feedback adjusts the fidelity penalty toward a shared accuracy-drop target. A privileged critic and all shadow computations are training-only. Deployment retains the frozen backbone and a lightweight actor that applies hard deletion and packed variable-length FlashAttention, converting token reduction into measured speedups. On ImageNet-1K with DeiT-Base, DORA reduces FLOPs by 38.4% relative to the uncompressed backbone within one percentage point of accuracy loss. Averaged across four ViT-type backbones at matched accuracy, DORA uses 13.2% fewer FLOPs and achieves 32.4% higher throughput than the corresponding per-backbone baseline means. Under zero-shot transfer to ImageNet-A, these gains widen to 20.3% and 45.6%, respectively.