Papers for

llm service providers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Large language models can be tricked to waste more computing power

FragToken: Amplifying LLM Inference Costs through Noncanonical Token Generation

Abstract: As large language model (LLM) inference becomes increasingly expensive, resource-consumption attacks pose a growing threat to model providers. Existing attacks typically amplify cost by inducing abnormally long or repetitive outputs on attacker-controlled or triggered requests, making them easier to detect and limiting their deployment-wide impact when benign traffic dominates. In this work, we uncover a previously overlooked token-level attack surface arising from the many-to-one mapping from token sequences to decoded text. Although standard LLMs predominantly generate the canonical token sequences induced by their tokenizers, the same text can also be represented by substantially longer non-canonical sequences. This representational flexibility exposes a new avenue for resource-consumption attacks: an attacker can train the model to favor such sequences, systematically increasing the number of autoregressive decoding steps without a proportional increase in visible response length. However, we empirically find that directly maximizing token fragmentation substantially degrades model utility, producing conspicuous answer-quality failures that undermine attack stealthiness. To address this challenge, we propose FragToken, a training-time framework that combines source-model self-distillation, capacity-aware filtering and budgeting, and BPE-Aligned Merging to induce fragmented generation under ordinary prompts while largely preserving model utility. We evaluate FragToken on four LLMs across three benchmarks. Across the four models, FragToken achieves a three-benchmark average token inflation ratio (TIR) ranging from 1.99 to 2.46, while causing only minor degradation in model utility. Our work reveals a covert LLM supply-chain threat that increases inference cost without requiring large volumes of attack requests while largely preserving utility.

Fri 25 SeptCryptography and Security
The gist
Large language models (LLMs) are costly to run, so attackers try to make them use more resources. The authors found a hidden trick where the model can produce the same visible text but using many more internal steps. This means the model works harder without showing longer outputs, making attacks harder to detect. They created FragToken, a method to train models to do this token trick while still giving useful answers, proving this hidden attack is real and effective.
Open → 2609.31552v1

Language models balance fast output and traceable text with new sampling method

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

Abstract: Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have shown that combining these two goals is highly nontrivial and can be potentially impossible. In this work, we develop a novel multi-draft speculative sampling algorithm based on Poisson processes that improves the frontier of this fundamental trade-off. The proposed algorithm has strong sampling efficiency on its own and, more interestingly, is naturally watermarkable: we can embed an unbiased watermark without degrading speculative acceptance. Moreover, our algorithm is based on an exact list-coupling-without-communication scheme, which yields a drafter invariance property that benefits both sampling and watermarking. It is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency, and we experimentally verify its strong performance in both aspects.

Fri 18 SeptCryptography and SecurityMachine Learning
The gist
Large language models can create text quickly and let people know that the text came from them. But doing both at the same time is hard. The authors designed a new way to pick words using a process based on random timing, which helps keep the text fast and easy to check. Their method also makes sure the text carries a special signature (a watermark) without slowing things down. They tested their idea and found it works well for both speed and watermark quality.
Open → 2609.21858v1

Provider-side attacks inflate large language model outputs and costs

The More It Says, the More You Pay: A Black-Box Audit of Provider-Side Token Inflation in LLM Services

Abstract: In pay-per-token LLM services, the more a model says, the more users pay. Dishonest providers can covertly manipulate generation to inflate output tokens while largely preserving task utility. We define such manipulation as a Provider-Side Token Inflation Attack (PTIA) and instantiate five representative attacks at the query, prompt, representation, and model levels of the provider-controlled pipeline. Our experiments show that each attack increases mean output length to more than 10.2x the clean baseline, demonstrating PTIA's financial appeal and feasibility at multiple stages of generation. Yet auditing PTIA from black-box responses is difficult for users. Our key observation is PTIA saturation: an initial attack sharply lengthens output, but further strengthening or composition has much less effect. We trace this saturation to stopping behavior: an initial PTIA sharply lowers the end-of-sequence token probability, whereas further intervention lowers it only marginally. Building on this insight, we design a lightweight single-probe audit that applies a controlled lengthening intervention. Under PTIA, the probe induces far fewer additional tokens than under normal service. The audit requires neither a trusted local reference model nor historical clean responses, and its separately issued original and probed requests resemble ordinary traffic, making evasion difficult. Across four open-weight models, it achieves an average detection rate of 85.1% with false-positive rates below 2%. Across 15 real LLM API services, the audit flags 7 for PTIA-consistent behavior.

Thu 17 SeptCryptography and Security
The gist
In some language model services, users pay based on how many words the model generates. The authors found that some providers might secretly increase the length of responses to make users pay more, without clearly changing helpfulness. They designed several ways this can be done and showed it’s practical. They also created a simple test that can detect when such manipulations happen, even without knowing the original clean response. When tested on multiple models and real services, this test flagged several services likely inflating output length.
Open → 2609.20370v1