Papers for

software engineers building ai assistants

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Exact method corrects bias in constrained text generation models

Twist, Don't Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models

Abstract: Constrained decoding for Masked Diffusion Language Models (MDLMs) aims to ensure that generated outputs satisfy a specified structure or syntax constraint. MDLMs generate outputs by repeatedly unmasking masked positions present in their current state. Recent strategies for constrained decoding constrain the model's per-step mean-field posterior (which factorizes over masked positions) by enforcing the desired constraint with an automaton. The resulting chain-structured factor graph allows exact constrained sampling via dynamic programming. However, despite each draw being exact and constraint-satisfying, we prove that their composition, in general, tilts away from the model's relative probabilities over valid trajectories, thus leading to trajectory bias. We derive an exact expression for this bias as a product of ratios measuring how valid continuation mass changes when the denoiser is reconditioned, and characterize when the bias vanishes. We then correct the bias by introducing TWISTER, the first automaton-twisted Sequential Monte Carlo decoder for MDLMs, using the step-exact decoder as the proposal. We show that for regular language constraints, the Feynman-Kac correction is exactly computable, with the twists obtained efficiently using quantities pre-computed for step-exact sampling. We prove that the resulting Feynman-Kac model targets the unbiased Doob h-transformed path law conditioned on constraint satisfaction.

Mon 28 SeptMachine LearningArtificial IntelligenceComputation and Language
The gist
Generating text with specific rules can be tricky for some AI models because the way they build sentences step-by-step can introduce subtle biases. The authors studied these biases in a type of masked diffusion language model, which fills in missing words repeatedly to create sentences that follow certain constraints. They found that previous methods caused a bias away from the model's true probabilities of valid word sequences. To fix this, the authors developed a new method called TWISTER that corrects the bias exactly, ensuring the generated text strictly follows the rules without skewing probabilities. This approach uses advanced sampling techniques to keep the outputs both valid and unbiased.
Open → 2609.35609v1

Large language models show big differences using similar tool interfaces

Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

Abstract: Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool definitions, and an agent that has truly learned a task should behave consistently across them. We show that current agents often do not, a phenomenon we term schema bias. To study this systematically, we introduce an executable transformation framework that rewrites a native tool schema using nine operators, including merging and splitting tools, altering how a single tool is expressed, and distributing one action across several dependent calls. The tasks, executable actions, and reachable states remain fixed, so any change in success is attributable to the interface alone. Evaluating eleven LLMs, including two closed models, on up to 32 schema variants, we ask how large schema bias is, how it manifests, whether the difficulty of a schema variant can be predicted without a full evaluation, and whether training removes it. We find that schema bias is substantial even for the newest models: success rates range from complete failure to 97% depending solely on the schema. To reliably estimate schema difficulty, it requires running a small sample of the target queries. Training repairs a schema variant only when that variant appears in the training data.

Mon 28 SeptArtificial Intelligence
The gist
When large language models (LLMs) use tools through different but equivalent interfaces, they often perform inconsistently. This problem, called schema bias, means that changing how a tool’s actions are presented can cause the same model to fail or succeed dramatically. The authors tested many LLMs with variations of tool interfaces that do the same thing but look different, finding large performance swings. They also showed that training a model helps only if the exact altered interface is in the training data. This reveals that LLMs do not fully understand the underlying tasks but rely heavily on how tool actions are shown.
Open → 2609.34971v1

Tessera cuts delays in memory-heavy AI model requests by up to 3.6 times

Tessera: Demand-Driven KV Cache Management for Retrieval-Augmented LLM Serving

Abstract: RAG and retrieval-based agent memory both inject retrieved content into LLM prompts, as document chunks and recalled memory records, respectively. The same content can recur across requests at different prompt positions or after different preceding contexts, preventing reuse through conventional prefix caching. Our characterization finds that records recurring outside the matching prefix account for over 70% of injected memory tokens in agent-memory workloads. Composable KV-reuse methods enable reuse in such cases, but online serving introduces a management problem: a recurring unit's KV states may not yet exist, may have been evicted, or may reside on another node. We present Tessera, a disaggregated serving system that makes retrieval the control plane for KV reuse. By exposing the context units needed before model execution, retrieval allows Tessera to combine current demand with retrieval history, KV residency, and generation load to coordinate cache management and request routing. Generation nodes concurrently prepare locally cached, remotely cached, and missing states, while retaining newly computed states off the request's critical path. Across RAG and agent-memory workloads, Tessera lowers mean TTFT by up to 3.6x over SGLang and LMCache with EPIC at matched request rates, and sustains low TTFT at rates where the baselines saturate, while matching the answer quality of the underlying composition policy.

Sat 26 SeptDistributed, Parallel, and Cluster Computing
The gist
Large language models (LLMs) often use extra information from documents or memory to answer questions better. However, reusing parts of previous computations to save time is tricky because this useful content can appear in different places or orders, making simple caching ineffective. The authors designed Tessera, a system that smartly manages caching by knowing what pieces of information are really needed upfront and sharing that knowledge across servers to prepare answers faster. Tessera speeds up response times significantly while keeping answer quality high, especially when many users ask questions at once.
Open → 2609.32999v1

Logical reasoning resides in a small brainlike part of language models

Logical subspace in LLMs

Abstract: Recent work has identified a human brain network specialized for abstract formal reasoning (Kean et al., 2025). Does the same hold true in language models? To answer this question, we introduce the minimal viable subspace (MVS) method, which searches for the lowest-rank activation subspace at a layer that preserves task performance when everything outside that subspace is ablated. Using MVS, we demonstrate low-rank subspaces supporting logical inference on Gemma and Qwen models. Furthermore, these subspaces exhibit a clear dissociation from model capacities on other tasks, such that retaining these late logic subspaces preserves inference while impairing factual knowledge, working memory, cognitive control, and arithmetic. Conversely, ablating them reduces logical inference accuracy to chance while largely sparing these other capacities. Our results suggest a functionally localizable core machinery for logic akin to that in the human brain.

Sat 26 SeptArtificial IntelligenceComputation and Language
The gist
Finding out whether language models have special parts for logical thinking like the human brain, the authors devised a way to pinpoint tiny areas in the model’s ‘brain’ that do the thinking. They found that these small parts handle logic tasks without helping with other abilities like memory or math. Removing these parts breaks logical thinking but leaves other skills intact. This shows that logical reasoning is localized in language models just as in human brains.
Open → 2609.32907v1

Language models lose structure when communicating complex expressions

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

Abstract: When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields a communication matrix whose marginals separate generation quality from extraction quality. Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60.4 points, and the best pair reaches 92.9% by combining different models on each end rather than the same model on both. Second, at least 73.6% of round-trip failures originate at generation, and difficulty is driven by tree structure (operator count, depth, right-branching) rather than model family. Third, the channel is trainable: ~3600 fine-tuning examples that share the evaluation's operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics. A disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains. Together these results identify tree-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language.

Fri 18 SeptArtificial Intelligence
The gist
Language models often translate complicated tree-like math expressions into text and then try to recover the original. This process is imperfect and means some details are lost or changed. The authors show that different models vary a lot in how well they can generate and understand these texts, and that training on similar examples helps. This loss limits how well models can share complex structured ideas through plain language.
Open → 2609.21509v1

Demonstration selection made efficient using state space models

Long-Context Demonstration Selection Using State Space Models

Abstract: We study the problem of demonstration selection, which involves selecting a subset of examples for prepending to a query to a language model. This problem is closely related to in-context learning and language model inference. Since the inference cost of a transformer model scales quadratically with sequence length, the selection problem becomes especially challenging in a long-context scenario. In this paper, we tackle this problem by building on state space models (SSMs), which require only linear inference time given the input. Our approach involves two algorithms. The first learns a small set of SSMs through distillation of a (trained) transformer model. We partition all the layers into consecutive groups. Then for each group, we estimate a separate state space model to replicate the input-output behavior within the adjacent layers. Second, we map the distilled model outputs to a small set of tokens, and apply these embeddings for demonstration selection in downstream applications. We perform extensive experiments in both synthetic and real-world datasets to validate our approach. We demonstrate that the distilled SSMs only incur an approximation error of less than $0.7\%$ relative to the true output. In downstream evaluation, we show that on several text classification and reasoning tasks, our approach reduces FLOPs by $14.2\times$ and improves accuracy by $6.48\%$ relative to baseline demonstration selection methods.

Tue 15 SeptMachine LearningComputation and Language
The gist
Choosing examples to show a language model before asking it a question can be slow and costly when the examples are long. The authors developed a method that uses simpler mathematical models called state space models to imitate parts of the original complex model more quickly. This lets them pick helpful examples faster and use less computing power without losing accuracy. Their tests showed this method works well on text classification and reasoning tasks, making the process both faster and more accurate.
Open → 2609.17888v1