ALICE estimates mutual information without retraining on new data
ALICE: In-context, Zero-shot, Mutual Information Estimation
Machine Learning
Summary
Mutual information measures how much two things are related, but calculating it usually needs lots of data and specific training for each case. The authors present ALICE, a model trained once on many simulated datasets so it can estimate mutual information for new, unseen data without extra training. ALICE works well even with small datasets and different types of data. This makes it easier to study relationships in complex fields like biology or neuroscience without preparing a new model each time.
What this means in practice
- •For bioinformatics teams: Conduct mutual information analysis on genetic and biological data without training bespoke models for each dataset.
- •For data engineers in neuroscience: Estimate relationships in neural datasets with fewer samples and no need for custom estimator retraining.
Authors
Giulio Franzese, Simone Rossi, Pietro Michiardi
Abstract
Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution under study. Current estimators are moreover tied to specific data types. These constraints limit their adoption in many applications where per-distribution training is impractical and sample sizes are small. We present ALICE, a foundation model that removes per-distribution training, while achieving competitive estimation accuracy. Trained exclusively on a broad family of synthetic distributions, ALICE acts as an in-context estimator of rectified-flow velocity fields: conditioned on samples of an unseen distribution, it estimates that distribution's velocity field without any explicit training. MI is then obtained through a fixed identity that integrates the squared difference between the joint and conditional fields. We validate ALICE on a standard, challenging benchmark and apply it in three domains, biology, genetics, and neuroscience, whose data the model has never seen. For the first time, we show that a single model closes the gap with neural estimators trained separately for each distribution, while natively supporting different data dimensionality and sample cardinality, enabling zero-shot MI analysis across scientific domains.