Papers for

cloud ai platform operators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

TACO optimizer slashes memory use for large language model tuning

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Abstract: Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO computes the exact steepest-descent direction under a dimension-normalized $1\to1$ operator norm by selecting the sign of the largest magnitude entry in each column of two-dimensional weight matrices. This retains first-order gradients while making optimizer state memory nearly negligible. Our practical TACO optimizer maintains only a small set of low precision gradient components per column, reducing persistent optimizer state by $174\times$ relative to AdamW8bit (from 27.7 GB to 0.16 GB) and peak training memory by $2.9\times$ (from 80.6 GB to 27.5 GB) on OPT-13B, while achieving comparable accuracy and runtime. TACO further enables full-parameter fine-tuning of 30-32B-parameter models on a single 80 GB H100 GPU across multiple model families and tasks.

Thu 1 OctMachine Learning
The gist
Fine-tuning large language models usually requires a lot of memory for storing optimizer information, which limits how big models can be on current GPUs. The authors propose TACO, a new optimizer that drastically reduces this memory need by storing only a small, simple piece of gradient information per column of model weights. This approach keeps accuracy and speed similar to current methods while allowing much larger models to be fine-tuned on a single GPU. TACO lets teams work with bigger language models without needing more powerful hardware upgrades.
Open → 2610.02199v1

Inference auction lets users bid for faster model responses

Inference Auctions

Abstract: When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction that allows users to bid for faster service. Our auction allocates priority in an economically efficient way without sacrificing latency, and we develop fast algorithms for implementing prices that incentivize truthful bidding. We also design an autobidding agent for our inference auction, where users specify an inference budget and the autobidder dynamically adjusts its bids over time to maximize user utility subject to the budget constraint. Experiments validate the practicality of our auction: it increases system welfare while maintaining the cache utilization and latency advantages of SGLang, a state-of-the-art inference serving framework.

Wed 30 SeptMachine LearningArtificial IntelligenceComputer Science and Game Theory
The gist
When lots of people ask a big AI model for answers at the same time, it can’t respond to everyone quickly. The authors created a way for people to bid money to get their answers faster, so the fastest answers go to those who value speed most. They built smart software that sets prices fairly and helps users bid within their budgets. Tests show this bidding system makes the whole AI-serving process work better without slowing it down.
Open → 2609.40070v1

Optimizer choice changes training speed but not data scaling rate

Optimizer-dependent training dynamics converge to the same one-third optimal data scaling

Abstract: Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size $D$. We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to $1/3$ across optimizers. In an online teacher-student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents $α_{r}$ and $α_{t}$. Under SGD, both are close to $1/3$, so the data exponent is also $1/3$ across different learning rates. Under Adam the two separate: $α_{r} \simeq 0.48$ but $α_{t} \simeq 0.08$. Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: $D$-independent for SGD but falls with $D$ for Adam. Yet tuned to that optimum, the loss returns to $D^{-1/3}$ for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, $2α_{r}+ α_{t} = 1$, which fixes the optimal data exponent at $1/3$. Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with $D^{-1/3}$ across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.

Tue 29 SeptMachine Learning
The gist
Training large language models involves reducing errors as they see more data. The authors show that different training methods (optimizers) affect how quickly the model improves during a single training run but do not change the fundamental relationship between data size and best achievable error rate. They find a consistent scaling pattern where the error drops in proportion to the dataset size raised to the power of one-third, regardless of the optimizer used. This reveals a shared underlying rule in how models learn from data, despite differences in optimization strategies.
Open → 2609.37745v1

Value-based token eviction improves cache use in advanced large language models

ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs

Abstract: Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturbance eviction under a max-entropy assumption about future attention. We evaluate under fixed cache budgets, with eviction at every block boundary during prefill and at every decoding step during generation. On RULER at a tight 2k token budget, ValueDiff retains 88--99\% of dense across seven sink-suppressed models (best on 6 out of 7). On LongBench at the 4k budget, ValueDiff averages 92\% retention across sink-suppressed models versus 83\% for the strongest prior baseline. On MATH-500, ValueDiff is the strongest non-dense method on every sink-suppressed model tested at the 25\% cache budget, outperforming prior methods by up to $\sim$20 points on gated-attention models. Across all three benchmarks, value geometry emerges as the more reliable query-invariant eviction signal for sink-suppressed models.

Sun 20 SeptMachine LearningArtificial Intelligence
The gist
Large language models often store past information temporarily to improve their responses, but they have limited memory space and need to decide which information to keep or remove. Existing methods rely on patterns that newer models with certain attention techniques don’t exhibit strongly, making those methods less effective. The authors noticed that paying attention to the variety in stored information values works better for these newer models. They developed ValueDiff, a method that measures how much each stored token’s information differs from the average and removes those that are least important, improving memory use and model accuracy.
Open → 2609.23314v1