Papers for

nlp system engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Improving reliability of language model activation explanations

Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation

Abstract: Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.

Sun 27 SeptComputation and Language
The gist
Language models have hidden layers that can be hard to understand because their explanations sometimes add incorrect or missing details. The authors propose a two-step way called AVPO that first tries to turn these hidden signals back into original text, then checks the accuracy of that text using a separate question-answering system. This method helps create clearer and more trustworthy explanations of what the model is thinking, reducing errors and making the recovered meaning more accurate.
Open → 2609.34033v1

Adaptive method improves large language model skill revision choices

StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills

Abstract: Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision forms applied under it. However, we observe that no single revision operator consistently performs best across tasks, and repeatedly applying an unsuitable operator can limit further improvement. We propose StraTune (strategy-guided skill tuning), which lets a frozen optimizer LLM choose the revision operator at every round from the optimization state, which is defined as the current execution feedback together with the recorded outcomes of earlier strategies and forms. Candidate skills from every revision operator pass one candidate evaluation, which screens for gains and regressions on a small sample set and validates them on a larger one, and every outcome is written back to the optimization state for later choices. Across four benchmarks and two LLM settings, StraTune outperforms all five baselines in most settings. Ablations attribute the gains to the adaptive choice of the revision operator, since fixed, random, scheduled, and bandit strategy choices all score lower, and skills learned with a small target LLM also improve a stronger one. Code and learned skills are available at https://github.com/seai-lab/StraTune.

Sat 26 SeptArtificial Intelligence
The gist
Large language models can learn new skills by revising how they perform tasks without changing their internal settings. The challenge is deciding the best way to make these revisions because no single method is best for all tasks. The authors created StraTune, a system that picks the best revision method at each step by learning from previous attempts and results. This adaptive approach helped the models improve skills better than using fixed revision methods across a set of tests.
Open → 2609.32886v1

Chain of thought reasoning steps often reflect the model’s actual calculations

Are Stated Reasoning Steps Causally Load-Bearing?

Abstract: Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning. Unlike previous causal audits, which measure degradation, our interventions carry a known predicted target. In this way, each patch should switch the answer to a specific counterfactual entity derivable by construction. Specifically, we use synthetic multi-hop lookup tasks (2-6 hops). We patch the residual stream at the token span where the model states each intermediate step with the corresponding activations from a counterfactual run. For Qwen3-4B, 76.9% +/- 2.8% of stated steps are causally load-bearing (CLB) at the most responsive mid-network layer (random-position null: 11.3%; patching the underlying prompt fact: 83%, so stated steps carry approximately 96% of the achievable effect). Moreover, the standard behavioral test on the same items yields 88.2%, which overstates causal faithfulness by 11.4 percentage points (item-matched; 111:14 discordant pairs, p < 1e-15) and, for the easiest items, by up to 20 percentage points. This gap also has a clear capability dimension. Qwen3-1.7B is far less causally faithful overall (54.8%), with its faithfulness collapsing as reasoning depth increases (68% at 2 hops to 30% at 6), while Qwen3-4B remains relatively flat. Although stated reasoning can be causally meaningful, standard behavioral tests tend to overestimate its causal faithfulness, particularly on easier examples where model reasoning appears most fluent.

Tue 22 SeptArtificial Intelligence
The gist
This paper looks at whether the reasoning steps AI models write down actually cause their answers or just look like reasoning. The authors study this by changing the model’s internal activity at the exact places where it states intermediate reasoning steps. They find that most of these stated steps truly affect the final answer, meaning they reflect the model’s real thought process. However, simpler behavioral tests tend to overestimate this effect, especially when the problem is easy or when the model is less capable.
Open → 2609.27038v1