Papers for

statistical modelers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Formula reveals loss patterns in shallow relu networks with bias

Population loss in shallow ReLU networks: Bias & families of critical points

Abstract: The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families of spurious minima described in past papers of Arjevani and the author extend to biased networks and that the loss is always strictly decreased when bias is added. The change in landscape geometry caused by adding bias appears to be relatively mild. Only the simplest examples are described in this paper where it is assumed that the number of inputs is equal to the number of neurons (this restriction is for reasons of length). A review of relevant previous results on unbiased networks is included. Aside from Gaussian statistics, the main mathematical tools and ideas come from analytic geometry (analytic and subanalytic sets, the Curve Selection Lemma).

Fri 25 SeptMachine Learning
The gist
The paper finds a formula that helps understand errors in simple neural networks that use ReLU and include a bias term. This extends earlier work that looked only at networks without bias. The formula involves a special mathematical function called Owen's T-function, which the authors explain and provide precise code for. They show that adding bias reduces the error and only slightly changes the network’s error landscape. The work focuses on the case where inputs and neurons have the same number, keeping the math simple.
Open → 2609.30661v1

Discrete diffusion model improves simulation of spin systems and sampling

Discrete Diffusion Models via Evolving Variational Autoregressive Networks

Abstract: Conventional score-based diffusion models learn scores without representing normalized densities, whereas tractable normalized models support both sampling and direct likelihood evaluation. A recent tensor-network approach provides such a representation but is largely restricted to low-dimensional lattices. Here we introduce a discrete diffusion model that parameterizes normalized probability distributions using variational autoregressive networks. Explicit Markov jump operators govern the forward noising and reverse denoising dynamics, extending discrete diffusion models with normalized distributions to spin systems on higher-dimensional lattices. We apply this framework to the two- and three-dimensional Ising models across ordered, critical, and disordered regimes, accurately computing thermodynamic quantities including free energy, energy, and magnetization. We further integrate the framework with Monte Carlo sampling, using adaptive diffusion steps to maintain high acceptance rates even at low temperatures while enhancing sample diversity. These results establish a neural-network framework for the discrete diffusion model with normalized probability distributions.

Wed 23 SeptMachine Learning
The gist
Models that describe complex systems often need to generate realistic examples and calculate how likely these examples are, but doing both is hard. The authors created a new method that represents probabilities exactly using special neural networks and simulates changes in systems like magnetic spins in multiple dimensions. Their method calculates important physical properties accurately and can work well with existing sampling techniques to produce diverse examples, even under difficult conditions. This approach helps better understand and simulate physical systems with complex interactions.
Open → 2609.27306v1

Copula operad framework links dependence and entropy additively

Copula Operad and Copula Entropy

Abstract: We construct a symmetric operad $\mathfrak{C}$ on the class of all multivariate copulas, with operadic composition given by Sklar substitution. This places the hierarchical combination of dependence structures into an algebraic framework. Restricting to absolutely continuous copulas whose densities lie in $L\log L$, we prove that copula entropy is additive under composition, making it an additive character on suitable finite-entropy suboperads. We exhibit explicit closed suboperads (bounded, $L^{p}$, boundary-growth) and show by counterexample that componentwise $L\log L$ does not imply closure under composition.

Thu 17 SeptInformation Theory
The gist
Understanding how variables relate to each other can be tricky, especially when combining multiple relationships at once. The authors created a mathematical framework that describes how these relationships, represented by something called copulas, combine together. They showed that a measure of uncertainty called copula entropy adds up neatly when combining copulas this way. However, not all types of copulas behave nicely under this combination, which they demonstrated with examples.
Open → 2609.20512v1

MaxEnt probability distributions on spheres inform models of cognition

Maximum Entropy Probability Distributions on Spheres with Fixed Mean Busemann Function and Holomorphic-Information-Geometric Model of Cognition

Abstract: In the first half of the paper we revisit the question regarding MaxEnt probability distributions on spheres. We derive families of MaxEnt distributions on spheres in real and complex vector spaces with fixed expected Busemann function (energy). As a particular case, we deduce sub-families on canonical energy levels where inverse temperature equals the dimension of the sphere. In the second part we focus on the information manifold of canonical MaxEnt distributions on the sphere in the complex vector space. This manifold is isomorphic to the Bergman ball. We introduce the reproducing kernel on this manifold and use the RKHS theory to elaborate the model of cognition. In particular, we state the principle of minimal cognitive effort in RKHS and demonstrate its dual relationship with the MaxEnt principle for probability distributions on the boundary sphere.

Thu 17 SeptInformation Theory
The gist
The paper studies the best ways to spread probabilities on spheres when we fix certain energy-like values. The authors find special families of these distributions in both real and complex spaces. Then they link these distributions to a geometric space called the Bergman ball, using advanced math tools called reproducing kernels. They propose a new way to think about cognition as minimizing effort in this geometric space, showing how it relates to the idea of maximum entropy for probabilities on the sphere.
Open → 2609.20410v1

Compressed subspaces speed up uncertainty analysis in large models

Compressed Active Subspaces for Scalable Bayesian Inference

Abstract: Active subspace methods provide a framework for quantifying predictive uncertainty in high-dimensional models by identifying and performing inference along parameter directions that have the greatest influence on the model output. However, the construction of active subspaces requires storing many full-dimensional model gradients, which becomes prohibitive as model size increases. We address this limitation by proposing Compressed Active Subspaces (CAS), a scalable approach that first maps the model parameters to a compressed space using a structured isometric embedding and then constructs the active subspace within this reduced parameterization. Our approach substantially reduces the memory required for active subspace construction and enables Bayesian inference for large models where standard active subspace methods become impractical. We demonstrate the scalability of CAS on neural networks of increasing size while maintaining predictive performance and robust uncertainty estimates.

Thu 17 SeptMachine LearningArtificial Intelligence
The gist
High-dimensional models often have many parameters, making it hard to understand which ones really affect the outcome. The authors propose a way to shrink these parameters into a smaller, compressed set without losing important information. This makes it easier and less memory-intensive to estimate uncertainty in predictions, especially for big models like neural networks. Their method still gives good predictions and reliable uncertainty estimates while being much more scalable.
Open → 2609.19539v1

Probability distribution evolution explained through least action principle

Physics of Information Geometry - Part I: Principle of Least Action on the Probability Simplex

Abstract: We develop a least-action framework for describing how a probability distribution can evolve from an equilibrium state to a prescribed nonequilibrium state under constrained incremental changes. Taking a Gibbs distribution as the equilibrium reference, the framework gives a direct physical meaning to the geometry of the probability simplex: distance from equilibrium corresponds to nonequilibrium free energy, while changes between successive distributions carry an informational kinetic cost. The Pythagorean structure of relative entropy then provides the central insight of the work. It shows that intermediate distributions chosen via sequential information projections can reduce the kinetic cost of large transitions and establishes an energy-conservation-like relation between the kinetic expenditure along a path and the free energy accumulated in reaching the target distribution. Motivated by this geometry, we construct a greedy least-action path through successive information projections, obtain a closed-form characterization of each projection through the Lambert W function, and establish a finite-step performance guarantee. We further show that state-dependent costs can be incorporated naturally by reshaping the underlying Gibbs reference, providing a thermodynamic interpretation of path penalties as modifications of the effective energy landscape. Together, these results provide a unified view of distributional evolution through least action, information geometry, and nonequilibrium thermodynamics.

Tue 8 SeptInformation Theory
The gist
The paper looks at how a probability distribution changes from a balanced state to an unbalanced one using a physics-inspired framework. It treats moving through possible probability states like a journey where distance measures how far from balance you are and movement costs information energy. By breaking big changes into smaller steps chosen carefully, less energy is used overall. The authors also connect this to thermodynamics, showing how the shape of the underlying probability landscape affects the costs to move through it.
Open → 2609.08285v1

Markov chains produce work only with non-Gaussian models in Bayesian inference

Thermodynamic Cyclic Processes with Markov Samplers in Bayesian Inference

Abstract: The concept of Markov chain Monte Carlo (MCMC) cycles, an analogy to cyclic processes in heat engines, is presented in order to examine Bayesian inference problems. In this effort, we develop adaptive ensemble schedulers that allow the tuning of external parameters of a Bayesian canonical ensemble during an MCMC run, realising the MCMC cycles in practice. We run these cycles on different statistical models. As a fundamental insight, we find (both theoretically and in practice) that such systems can produce a non-zero net work output if and only if the considered model is non-Gaussian. As such, they may serve as a measure of non-Gaussianity in Bayesian inference, which we test on an example from supernova cosmology.

Mon 7 SeptArtificial Intelligence
The gist
The paper looks at Bayesian inference—the process computers use to make decisions based on data—and draws an analogy to heat engines that do cycles. The authors develop a way to run Markov chain Monte Carlo cycles by tuning parameters during sampling. They find that these cycles can produce a measurable effect, or "work," only when the statistical model is non-Gaussian, meaning it doesn't follow the simple bell curve. This discovery offers a way to tell if data or models are non-Gaussian, tested on a cosmology example with supernova data.
Open → 2609.07660v1