Language model cells struggle to share and reuse communication codes
Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells
Artificial Intelligence
Summary
This paper studies how different groups of language-model components communicate using hidden 'languages' or codes. The authors find that independently trained groups develop different communication systems that are mostly incompatible, making it hard for them to understand each other without extra adaptation. They also show that trying to use a common communication method learned from a global model can actually hurt learning in new models. This research focuses on understanding the challenges of reusing and transferring hidden communication in AI models.
What this means in practice
- •For machine learning engineers: Identify when transferring communication protocols between independently trained model components may require adaptation to avoid negative performance impacts.
- •For ai system integrators: Plan effective integration of separately trained language-model modules by testing compatibility of their latent communication codes.
Tested on simulated data.
Authors
Narcis Marincat
Abstract
In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whether independently trained societies share one packet language, where strict zero-shot transfer fails, and whether inherited interface state helps or harms later learning. First, a leakage-controlled causal interoperability audit over all 30 ordered pairs of six independently trained restricted societies -- under sealed held-out structure and a preregistered raw/orthogonal/linear/nonlinear alignment ladder -- shows the six semantically similar interfaces do not form one raw language: one same-initialization pair is exactly interoperable in both directions, a second shows asymmetric partial compatibility, and all 26 cross-initialization directions fail every frozen alignment rung. Second, within the tested decomposition and a single sealed source formulation, a source-span control localizes strict zero-shot failure to interpretation and execution of the new operator instructions. Third, in a matched adaptation factorial, the globally trained communication interface acts as a severe negative-transfer prior: reinitializing only the packet reader, writer, and mouth raises final depth-three accuracy from 0.169 to 0.857. Fourth, across two restricted checkpoints and two independently frozen target streams each, inherited interfaces never exceeded fresh-interface controls by the preregistered 0.10 margin. All primary conclusions are bounded to a near-transfer 17-state setting; the negative-transfer factorial concerns one globally visible parent-cohort checkpoint, while an appendix adds a post hoc tagged-global twin case study.