Papers for

llm developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Parametric memory influences large language models reasoning performance

MemoReason: Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMs

Abstract: Large Language Models (LLMs) perform well on reasoning benchmarks, but it remains unclear whether this reflects genuine contextual reasoning or reliance on facts memorized in their parameters. We investigate this by distinguishing two possibilities: a broad \textit{memorization bias}, where familiar content improves reasoning performance, and the \textit{Strong Parametric Shortcut Hypothesis}, where models skip reasoning entirely and recall stored answers. To test these effects, we introduce \textbf{MemoReason}, a human-curated benchmark that pairs factual reasoning tasks with structurally identical \fictitiousterm{} versions where real entities like people, companies, or dates are systematically replaced by \fictitiousterm{} ones of the same type. This \scorerevision{preserves task structure and specified reasoning operations} while varying the familiarity of the context, allowing controlled measurement of how the parametric memory affects reasoning. \revision{Our evaluation of recent LLMs reveals consistent and statistically significant performance drops of up to 15.7\% in the fictitious setting, demonstrating a clear memorization bias.} However, a targeted analysis of \revision{questions failed in the fictitious setting} shows that models rarely respond with the corresponding factual answer, indicating that direct parametric shortcuts are not the dominant failure mode. These findings suggest that parametric memory influences reasoning through mechanisms more complex than simple factual recall. \textbf{MemoReason} provides a controlled framework for studying these mechanisms and for extending paired factual-fictitious{} evaluation to broader reasoning settings.

Mon 28 SeptComputation and Language
The gist
Large language models (LLMs) do well on reasoning tests, but it was unclear if they actually think through problems or just remember facts they learned before. The authors created a test called MemoReason, which swaps real facts for made-up ones while keeping the questions the same. They found that the models perform worse with made-up facts, showing they rely on memory, but they don’t just spit out memorized answers without reasoning. This means the models’ memory affects reasoning in more complex ways than just recalling facts.
Open → 2609.35312v1

Directional study reveals hidden risks in AI model safety training

See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs

Abstract: Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens. Grassmannian projections show that, in most settings, harmful-safe gap widens mainly because safe-gradient overlap declines. Leveraging these insights, we introduce a parameter-level Geometric Mitigation Framework that orthogonally projects empirical harmful gradient subspace out of parameter updates. On Qwen2.5-14B-IT, our defense suppresses free-generation EM by up to 80.0%; across the other three of four open-weight instruction-based model families (3B--20B), where single-layer behavioral EM is already near zero, teacher-forced evaluation shows same harmful subspace controls the conditional support of frozen EM responses. Crucially, these diagnostics unmask the illusion of behavioral safety: the same subspace remains measurable and steerable in models where behavioral EM is near zero. Code: https://github.com/WeiqiaoQUE/mechanistic-emergent-misalignment.

Mon 28 SeptMachine LearningComputation and Language
The gist
Large language models trained for safety can still unexpectedly fail when adapting to new topics. The authors studied how these failures develop during training by looking closely at the model’s learning directions and found that certain key words cause sharp risk patterns. They created a method that removes risky learning directions to keep the model safer without losing usefulness. Their tests show this method can reduce unexpected safety failures by a large amount while revealing hidden risk patterns even when the model looks safe.
Open → 2609.34970v1

LLM reasoning accuracy varies across sampled geometry problem answers

Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence

Abstract: Repeated sampling can reveal a correct numerical answer without yielding either a reliable system output or a supported derivation. We present Coverage, Realization, and Validity Evidence (CRV), an evaluation protocol for sampled large language model (LLM) reasoning over formal geometry states. Coverage is answer availability, realization is readout accuracy on the frozen candidate pool, and validity evidence is a label-blinded critic judgment of derivational support rather than a proof certificate. CRV freezes each candidate pool before comparing readouts and analyzes covered failures by correct-answer multiplicity and within-problem discrimination. On HardShift441, a 441-problem set for which a reference solver leaves 406 problems unsolved, a LoRA-adapted Qwen2.5-7B generator obtains 24.2% average single-sample accuracy and 68.9% pass@16, whereas verifier-weighted self-consistency (WSC) reaches 38.0%. Readout accuracy is particularly low when the correct answer occurs only once or twice in the pool. In a separate constructed audit of 195 covered problems, the critic labels 12 correct-answer representatives as supported, 181 as refuted, and two as uncertain. These results show that coverage, realization, and validity evidence from the critic are distinct quantities and should be reported separately.

Sat 26 SeptArtificial Intelligence
The gist
When large language models (LLMs) try to solve formal geometry problems, taking multiple guesses can find the right number but not always a trustworthy explanation or confirmed solution. The authors introduce a way to measure three things: how often the right answer appears (coverage), how correctly the model reads those answers (realization), and how well an expert judge thinks the reasoning supports the answer (validity evidence). They tested this on a hard set of 441 problems and found that simply having the right answer in samples doesn’t guarantee accurate or well-supported solutions. These three measures capture different aspects of quality and should be reported separately.
Open → 2609.32924v1

Adaptive method improves large language model skill revision choices

StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills

Abstract: Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision forms applied under it. However, we observe that no single revision operator consistently performs best across tasks, and repeatedly applying an unsuitable operator can limit further improvement. We propose StraTune (strategy-guided skill tuning), which lets a frozen optimizer LLM choose the revision operator at every round from the optimization state, which is defined as the current execution feedback together with the recorded outcomes of earlier strategies and forms. Candidate skills from every revision operator pass one candidate evaluation, which screens for gains and regressions on a small sample set and validates them on a larger one, and every outcome is written back to the optimization state for later choices. Across four benchmarks and two LLM settings, StraTune outperforms all five baselines in most settings. Ablations attribute the gains to the adaptive choice of the revision operator, since fixed, random, scheduled, and bandit strategy choices all score lower, and skills learned with a small target LLM also improve a stronger one. Code and learned skills are available at https://github.com/seai-lab/StraTune.

Sat 26 SeptArtificial Intelligence
The gist
Large language models can learn new skills by revising how they perform tasks without changing their internal settings. The challenge is deciding the best way to make these revisions because no single method is best for all tasks. The authors created StraTune, a system that picks the best revision method at each step by learning from previous attempts and results. This adaptive approach helped the models improve skills better than using fixed revision methods across a set of tests.
Open → 2609.32886v1

Post-training increases confident wrong answers in language models

The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models

Abstract: Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data, reasoning errors, or stochastic decoding, we uncover that post-training alignment itself is a primary driver of these errors, a phenomenon we call the \textbf{Alignment Paradox}. Across five model families evaluated on factual benchmarks, unaligned base models produce few high-confidence errors on long-tail factual queries, whereas instruction-tuned models multiply high-confidence errors ($p \ge 0.95$) by more than an order of magnitude (10$\times$ to 35$\times$). Layer-wise probing with the Logit Lens reveals that this overconfidence emerges in late layers, where wrong-answer margins expand past 4.0 points after remaining near zero across early and intermediate layers. These findings motivate limiting margin growth during post-training. We implement this principle through an entropy-dependent margin bound in direct preference optimization (DPO). In multi-epoch experiments with Mistral-7B, the bounded objective reduces high-confidence errors by up to 35.3\% relative to standard DPO while maintaining performance on evaluated general reasoning benchmarks. These results show that bounded margins mitigate confident hallucinations during post-training.

Sat 26 SeptComputation and LanguageArtificial Intelligence
The gist
Large language models sometimes give wrong answers that sound very sure, which makes it hard to tell when they are mistaken. This paper finds that fine-tuning these models to follow instructions actually makes them more confident in their incorrect answers, especially for rare facts. The authors identify that this overconfidence appears in the last layers of the model after fine-tuning. They propose a way to limit this overconfidence during post-training, which reduces confident errors while keeping the model’s reasoning abilities intact.
Open → 2609.32617v1

Hard prompts are limited and hard to optimize for large language models

On the Capability and Limitation of Hard Prompt

Abstract: Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical hard or discrete prompts is largely open. In this paper, we try to fill this gap either by providing a complete solution or by making substantial progress on the three core theoretical questions regarding hard prompts. First, we show that determining the existence of a hard prompt for a transformer to solve a downstream task is NP-complete and that finding an optimal hard prompt is NP-hard, which is the first computational complexity result for hard prompting, as far as we know. Second, we show that, unlike soft or continuous prompts, hard prompts have essential limitations: hard prompts are not complete; short hard prompts do not significantly enhance the ability of transformers; and long hard prompts exhibit the "prompt dominating answer phenomenon," meaning that, with high probability, the same answer is given for all queries of the same length. On the other hand, linear hard prompts do not have the limitations of short or long prompts. Third, we provide a tight bound on the size of the task in terms of the prompt length for the performance of prompts on the finite task to generalize to the entire data distribution, leading to a necessary and sufficient condition for generalizability. This is the first result on generalization for prompting, as far as we know. Our findings not only offer the first theoretical insights into hard prompts but also provide provably reliable practical guidance for real-world LLM usage.

Sat 26 SeptMachine Learning
The gist
Hard prompts are exact word sequences used to guide language models instead of adjusting their internal settings. The authors show it’s very difficult to find good hard prompts because the problem is computationally complex. They found that short or long hard prompts have intrinsic problems: short ones don’t help much, and long ones often force the model to give the same answer regardless of the question. They also provide rules about when hard prompts can work well across different tasks, offering the first theoretical understanding of how and when hard prompts generalize.
Open → 2609.32302v1

User simulator improves dialogue realism and persuasion accuracy

Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

Abstract: Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.

Wed 23 SeptArtificial Intelligence
The gist
Simulating user conversations is important to make AI systems that chat feel real and helpful. The authors created a new simulator called TRACER that better mimics how users change their goals during a conversation and how conversations usually end. TRACER learns from real dialogues and then improves by practicing many-turn conversations to match real behaviors closely. It outperforms previous simulators, produces realistic chat flows, and helps test AI models for both conversation quality and persuasion success.
Open → 2609.28690v1

Language model outputs traced to key internal components

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

Abstract: Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We supervise training over all sparsities simultaneously by randomising $k$ over training, resulting in a learned ordering of components by attribution score. MAttr achieves number 1 on the official leaderboard of the Mechanistic Interpretability Benchmark (Mueller et al., 2025); our method identifies sparse and task-transferrable circuits across varying circuit bases. As a practical application, we show that MAttr can be trained with reinforcement learning to identify weight changes responsible for downstream behaviours in LLM finetuning. We train MAttr on refusal judge scores and find that restoring $1\%$ of Llama 3.1 8B Instruct's weights to their base model state is sufficient to remove refusals while maintaining capabilities. We view MAttr as a successful formulation of interpretability into a learnable objective that we can tackle with gradient descent, and encourage future work along these lines.

Tue 22 SeptComputation and LanguageMachine Learning
The gist
It is hard to figure out which parts inside large language models actually make their answers. The authors propose a new way called Matryoshka Attribution that learns to find nested groups of important components by minimizing errors. Their method can rank parts by how much they contribute and found small groups that explain model behavior well. They also showed it can identify the weight changes causing an updated model to refuse certain requests, while keeping other abilities intact.
Open → 2609.25518v1

Attention patterns reveal hallucination in large language models

Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing

Abstract: In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman-Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinated responses. We evaluate our approach extensively across several LLMs and established benchmarks. Empirical results demonstrate that our proposed single-pass approach provides consistent improvements over existing attention-based and multi-response baselines across two hallucination-detection benchmarks, while achieving competitive performance across diverse LLM architectures. Further analysis reveals that impaired context sharing among tokens during causal generation is strongly associated with hallucination occurrences in LLMs. In particular, hallucinated responses are consistently characterized by an over-reliance on self-attention, diffused context retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.

Thu 17 SeptArtificial IntelligenceComputation and LanguageMachine Learning
The gist
Sometimes, large language models make mistakes called hallucinations, where they give wrong or made-up answers. The authors studied how these models share information inside themselves by looking at how their attention mechanisms connect words. They found that hallucinated answers show a breakdown in how the model passes context between words, especially relying too much on certain self-focused attention. This approach helps spot when the model is likely hallucinating by analyzing these connection patterns.
Open → 2609.21096v1

Dual-axis optimization method improves learning for large language model agents

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

Abstract: Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.

Thu 17 SeptArtificial Intelligence
The gist
Training large language model agents to perform tasks based on feedback is complicated because it involves two steps: figuring out which parts of their actions led to feedback inside a single attempt, and combining results across many attempts. The authors introduce a new way called BATON that handles these two steps separately but together. One part uses Bayesian methods to better assign credit to actions inside a single attempt, while the other balances how many complete attempts are considered equally during training. Tests show that using both techniques together helps agents learn better on different tasks and models.
Open → 2609.19830v1

Suan improves safety and usefulness in large language models

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

Abstract: Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynamics. Extensive evaluations across a diverse suite of competitive baselines and benchmarks demonstrate that Suan achieves superior safety alignment while fully preserving response utility.

Tue 8 SeptMachine LearningArtificial Intelligence
The gist
Large language models can sometimes give unsafe or unhelpful answers, and making them safer without reducing quality is hard. The authors present Suan, a new way to train these models that directly changes how the model learns from preferences. This new method helps keep the model both safe and useful better than older techniques. They tested Suan extensively and found it outperforms other approaches on diverse safety benchmarks.
Open → 2609.08634v1