Improving reliability of language model activation explanations
Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation
Computation and Language
Summary
Language models have hidden layers that can be hard to understand because their explanations sometimes add incorrect or missing details. The authors propose a two-step way called AVPO that first tries to turn these hidden signals back into original text, then checks the accuracy of that text using a separate question-answering system. This method helps create clearer and more trustworthy explanations of what the model is thinking, reducing errors and making the recovered meaning more accurate.
What this means in practice
- •For ai model developers: Create interpretable tools that generate more trustworthy explanations of language model internal states for debugging and model understanding.
- •For nlp system engineers: Integrate more reliable activation verbalization methods to improve monitoring and evaluation of text generation models in production systems.
Authors
Haiyan Zhao, Zirui Hei, Wei Shi, Huiqi Deng, Na Zou, Mengnan Du
Abstract
Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to trust and use reliably in practice. To this end, we introduce AVPO, a two-stage framework that first reconstructs source text from a hidden activation and then evaluates the resulting text with a separate frozen question-answering model, yielding an explicit and inspectable intermediate readout. We further optimize the inverter with direct preference optimization (DPO), using rewards that capture both semantic recoverability and lexical fidelity. Across six text families, AVPO improves gist- and detail-level information recovery over the strongest baseline by up to 17.1 and 9.3 percentage points, respectively. Crucially, the gains arise from preference optimization rather than fine-tuning on selected reconstructions alone, enabling compact cross-model inverters to surpass donor-matched question-conditioned verbalizers while improving both semantic recoverability and lexical fidelity. Moreover, out-of-distribution case study shows that AVPO better recovers high-level semantics while fabricating fewer details.