Papers for
statistical modelers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Formula reveals loss patterns in shallow relu networks with bias
Population loss in shallow ReLU networks: Bias & families of critical points
Abstract: The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families of spurious minima described in past papers of Arjevani and the author extend to biased networks and that the loss is always strictly decreased when bias is added. The change in landscape geometry caused by adding bias appears to be relatively mild. Only the simplest examples are described in this paper where it is assumed that the number of inputs is equal to the number of neurons (this restriction is for reasons of length). A review of relevant previous results on unbiased networks is included. Aside from Gaussian statistics, the main mathematical tools and ideas come from analytic geometry (analytic and subanalytic sets, the Curve Selection Lemma).
Discrete diffusion model improves simulation of spin systems and sampling
Discrete Diffusion Models via Evolving Variational Autoregressive Networks
Abstract: Conventional score-based diffusion models learn scores without representing normalized densities, whereas tractable normalized models support both sampling and direct likelihood evaluation. A recent tensor-network approach provides such a representation but is largely restricted to low-dimensional lattices. Here we introduce a discrete diffusion model that parameterizes normalized probability distributions using variational autoregressive networks. Explicit Markov jump operators govern the forward noising and reverse denoising dynamics, extending discrete diffusion models with normalized distributions to spin systems on higher-dimensional lattices. We apply this framework to the two- and three-dimensional Ising models across ordered, critical, and disordered regimes, accurately computing thermodynamic quantities including free energy, energy, and magnetization. We further integrate the framework with Monte Carlo sampling, using adaptive diffusion steps to maintain high acceptance rates even at low temperatures while enhancing sample diversity. These results establish a neural-network framework for the discrete diffusion model with normalized probability distributions.
Copula operad framework links dependence and entropy additively
Copula Operad and Copula Entropy
Abstract: We construct a symmetric operad $\mathfrak{C}$ on the class of all multivariate copulas, with operadic composition given by Sklar substitution. This places the hierarchical combination of dependence structures into an algebraic framework. Restricting to absolutely continuous copulas whose densities lie in $L\log L$, we prove that copula entropy is additive under composition, making it an additive character on suitable finite-entropy suboperads. We exhibit explicit closed suboperads (bounded, $L^{p}$, boundary-growth) and show by counterexample that componentwise $L\log L$ does not imply closure under composition.
MaxEnt probability distributions on spheres inform models of cognition
Maximum Entropy Probability Distributions on Spheres with Fixed Mean Busemann Function and Holomorphic-Information-Geometric Model of Cognition
Abstract: In the first half of the paper we revisit the question regarding MaxEnt probability distributions on spheres. We derive families of MaxEnt distributions on spheres in real and complex vector spaces with fixed expected Busemann function (energy). As a particular case, we deduce sub-families on canonical energy levels where inverse temperature equals the dimension of the sphere. In the second part we focus on the information manifold of canonical MaxEnt distributions on the sphere in the complex vector space. This manifold is isomorphic to the Bergman ball. We introduce the reproducing kernel on this manifold and use the RKHS theory to elaborate the model of cognition. In particular, we state the principle of minimal cognitive effort in RKHS and demonstrate its dual relationship with the MaxEnt principle for probability distributions on the boundary sphere.
Compressed subspaces speed up uncertainty analysis in large models
Compressed Active Subspaces for Scalable Bayesian Inference
Abstract: Active subspace methods provide a framework for quantifying predictive uncertainty in high-dimensional models by identifying and performing inference along parameter directions that have the greatest influence on the model output. However, the construction of active subspaces requires storing many full-dimensional model gradients, which becomes prohibitive as model size increases. We address this limitation by proposing Compressed Active Subspaces (CAS), a scalable approach that first maps the model parameters to a compressed space using a structured isometric embedding and then constructs the active subspace within this reduced parameterization. Our approach substantially reduces the memory required for active subspace construction and enables Bayesian inference for large models where standard active subspace methods become impractical. We demonstrate the scalability of CAS on neural networks of increasing size while maintaining predictive performance and robust uncertainty estimates.
Probability distribution evolution explained through least action principle
Physics of Information Geometry - Part I: Principle of Least Action on the Probability Simplex
Abstract: We develop a least-action framework for describing how a probability distribution can evolve from an equilibrium state to a prescribed nonequilibrium state under constrained incremental changes. Taking a Gibbs distribution as the equilibrium reference, the framework gives a direct physical meaning to the geometry of the probability simplex: distance from equilibrium corresponds to nonequilibrium free energy, while changes between successive distributions carry an informational kinetic cost. The Pythagorean structure of relative entropy then provides the central insight of the work. It shows that intermediate distributions chosen via sequential information projections can reduce the kinetic cost of large transitions and establishes an energy-conservation-like relation between the kinetic expenditure along a path and the free energy accumulated in reaching the target distribution. Motivated by this geometry, we construct a greedy least-action path through successive information projections, obtain a closed-form characterization of each projection through the Lambert W function, and establish a finite-step performance guarantee. We further show that state-dependent costs can be incorporated naturally by reshaping the underlying Gibbs reference, providing a thermodynamic interpretation of path penalties as modifications of the effective energy landscape. Together, these results provide a unified view of distributional evolution through least action, information geometry, and nonequilibrium thermodynamics.
Markov chains produce work only with non-Gaussian models in Bayesian inference
Thermodynamic Cyclic Processes with Markov Samplers in Bayesian Inference
Abstract: The concept of Markov chain Monte Carlo (MCMC) cycles, an analogy to cyclic processes in heat engines, is presented in order to examine Bayesian inference problems. In this effort, we develop adaptive ensemble schedulers that allow the tuning of external parameters of a Bayesian canonical ensemble during an MCMC run, realising the MCMC cycles in practice. We run these cycles on different statistical models. As a fundamental insight, we find (both theoretically and in practice) that such systems can produce a non-zero net work output if and only if the considered model is non-Gaussian. As such, they may serve as a measure of non-Gaussianity in Bayesian inference, which we test on an example from supernova cosmology.