Small transformers learn to track hidden causes across contexts

Small transformers track Bayesian evidence for latent common causes via a context-invariant mechanism

Machine Learning

Summary

The paper studies how small transformer models can learn to reason about hidden common causes behind data, using ideas from Bayesian reasoning. The authors focus on how these models figure out these causes across different situations, without being tied to one specific example. They show that the models can represent and accumulate evidence for these hidden causes in a way that works in new contexts. This helps understand how AI can generalize about underlying reasons for what it sees, like deducing a shared source behind different observations.

What this means in practice

Authors

Amir Mohammadpour, Michael Franke

Abstract

We present an in-depth investigation of how a form of Bayesian reasoning about common causes can emerge as a cross-contextual generalization in small, tractable transformers. Incrementing on recent work, our set-up (i) disentangles causal mechanisms in the model from the causal structure of the true data-generating process, (ii) orients more towards natural language prediction by considering inference of latent common causes, and (iii) considers whether and how Bayesian evidence accumulation for latent common causes can be implemented in representations and mechanisms that allow for cross-context generalization to novel test cases.