Papers for

ai model deployers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SparseOPD improves efficiency of on-policy distillation with selective corrections

Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation

Abstract: On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.

Mon 28 SeptComputation and LanguageMachine Learning
The gist
When teaching AI models to generate better answers, a method called on-policy distillation uses corrections from a teacher model on all possible words. But checking every word takes a lot of memory for long answers. The researchers introduced SparseOPD, a smarter way that first looks at all corrections to pick only the important words before updating the model. This reduces memory use dramatically while keeping or improving the quality of the AI’s answers.
Open → 2609.34386v1

MpFA boosts long-context AI speed on NVIDIA Blackwell GPUs

MpFA: Hardware-Efficient Train-Free QK4V8 FlashAttention Kernels on Blackwell GPUs

Abstract: Long-context LLM inference pushes modern GPU serving stacks into an attention-bound regime, where both compute and memory are dominated by the softmax-GEMM pipeline. On NVIDIA Blackwell GPUs, FP4 Tensor Cores offer high matmul throughput, but we find that fully FP4 attention often fails to translate this throughput into end-to-end speedups due to non-matmul costs: online quantization after softmax, tensor/shared-memory data movement, and contention on the softmax path. We present MpFA, a training-free FlashAttention kernel optimized for Blackwell. Guided by hardware characterization, MpFA uses mixed precision: NVFP4 for QK and FP8 for PV (QK4PV8). This preserves low-bit QK throughput while avoiding the conversion and scaling overheads of FP4 PV. To recover accuracy without further stressing the softmax pipeline, MpFA introduces rank-one smoothing compensation implemented as an additional Tensor Core MMA. MpFA further improves performance with a fine-grained asynchronous pipeline, tensor-memory reuse, and adaptive parallel partitioning across prefill and decode. On an NVIDIA B200 and across 16K-128K contexts, MpFA improves prefill throughput over state-of-the-art BF16/FP8 baselines and increases end-to-end output throughput by 2.81$\times$ over BF16 FA4 across Llama-3.1-8B and Qwen3-14B. Across five benchmark suites and two models, rank-one compensation recovers 62.5% of the accuracy loss with about 2.0% kernel overhead.

Sun 27 SeptDistributed, Parallel, and Cluster Computing
The gist
Long-context AI models require a lot of memory and computing power, slowing down their operations on GPUs. The authors introduced MpFA, a new method that mixes different precisions of math operations to speed up a key part of these models without needing retraining. They also found a way to fix accuracy lost from this change using a clever smoothing step. Overall, MpFA makes AI models run faster when handling very long inputs on certain NVIDIA GPUs.
Open → 2609.33135v1