An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers
2026-07-20 • Machine Learning
Machine Learning
AI summaryⓘ
The authors develop a mathematical framework to analyze how different positions within causal residual Transformers affect the model's behavior. They separate general results that hold unconditionally from conditions that depend on specific boundary shapes. Their work includes a detailed decomposition of how influence flows through layers and attention mechanisms, showing that certain factors can increase sensitivity to early positions but do not alone create a U-shaped influence pattern. The paper also proposes diagnostic tools to assess and control positional influence during training, supported by simulations that validate each tool's effectiveness.
Causal residual TransformerAdjoint sensitivityVolterra attentionGradient flowPositional influenceResidual connectionsEnergy densityLayer controlBoundary effectsInfluence diagnostics
Authors
Cheng Huan, Hongwei Yuan
Abstract
We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions. The principal unconditional theorem is the residual-to-depth-flow estimate for layer controls converging in $L^1$, complemented by a finite-token-to-Volterra attention estimate that explicitly controls the first cells near the causal endpoint. We define a normalized adjoint-energy influence density and derive its exact evolution along full-batch gradient flow. The adjoint admits an exact generator-term decomposition into residual transmission, nonlocal Volterra, and local channels, including all covariance cross terms. Causal masking can amplify early-position sensitivity and residual identity paths can transmit a right-localized terminal bias, but neither mechanism alone forces a U-shaped profile. We therefore state boundary advantages under independently checkable energy, correlation, and local-channel bounds; these conditions are sufficient rather than necessary. Finite-token influence balancing, positional reweighting, and task-aligned observability are presented as diagnostics or regularizers with explicit differentiation requirements, computational costs, and limitations. Controlled simulations illustrate that each intervention controls its designated surrogate, while observability balance or outer-loop reweighting need not monotonically reduce the influence-based Lost-in-the-Middle diagnostic.