Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner

2026-08-17Machine Learning

Machine LearningArtificial Intelligence
AI summary

The authors studied how a complex neural network learns to solve reasoning tasks, comparing its behavior and what can be decoded from its internal states during training. They found that the network’s ability to show correct answers and the clues inside its hidden layers develop on very different timescales and don’t neatly align. Probing the internal states before the network behaves correctly showed early hints of knowledge, but this didn’t perfectly match its actual performance. The authors conclude that just observing behavior or decoding internal signals isn’t enough to understand the learning process fully; direct intervention is needed to pinpoint what the network has actually learned.

recurrent neural networksrelational reasoningbehavioral thresholdshidden state probingtraining dynamicslinear probescausal interventionsymbolic reasoningmodel interpretabilitylearning development
Authors
Simon Lam-Muir
Abstract
Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a hidden state. These quantities need not identify the same event. We study a 30M-parameter recurrent-depth relational reasoner in a closed, oracle-defined world, using dense behavioural trajectories, two training surfaces, preregistered pre-arrival hidden-state probes, prospectively checked evaluability, and explicit untrained and negative controls, holding the training-time and inference-time axes separate throughout. Behaviour first: under one frozen acquisition criterion, three-hop competence cost 70 logical epochs on the symbolic surface and 13,055 on the verbal surface, a 186.5-fold contrast, after which verbal four-hop competence cleared in 8 logical epochs. Across the 13,055-epoch grind, four-hop held-out behaviour never exceeded 3/40 and ended at 0/40. Internal measurement next: on the verbal surface a linear probe recovered future-answer identity before behavioural arrival at 0.056159 against uniform chance 0.025, an untrained control of 0.024758 and a population frequency baseline of 0.048309 (p = 0.012987; 16/40 answer classes contributing). Analogous pre-arrival accessibility survived the surface change, reaching 0.1020 against a zero-step control of 0.0460 (p = 0.000999) at the upstream structural position and 0.0618 at the readout comparator (p = 0.004), with 21/40 classes contributing. Finally, the natural attempt to track that accessibility across training was not cleanly evaluable: probe eligibility is defined by behavioural arrival, so the measured population changes with the measurand. Behavioural competence, internal accessibility, and training-time development are distinct observables, and neither behaviour nor decoder accessibility identifies the computation training acquired; causal intervention is the necessary next step.