Formula reveals loss patterns in shallow relu networks with bias
Population loss in shallow ReLU networks: Bias & families of critical points
Machine Learning
Summary
The paper finds a formula that helps understand errors in simple neural networks that use ReLU and include a bias term. This extends earlier work that looked only at networks without bias. The formula involves a special mathematical function called Owen's T-function, which the authors explain and provide precise code for. They show that adding bias reduces the error and only slightly changes the network’s error landscape. The work focuses on the case where inputs and neurons have the same number, keeping the math simple.
What this means in practice
- •For machine learning engineers: Evaluate how adding bias affects training loss landscapes in shallow ReLU models to improve optimization strategies.
- •For statistical modelers: Use the loss formula to analyze error behavior under Gaussian assumptions for kernel-based shallow networks with bias.
Tested on simulated data.
Authors
Michael Field
Abstract
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families of spurious minima described in past papers of Arjevani and the author extend to biased networks and that the loss is always strictly decreased when bias is added. The change in landscape geometry caused by adding bias appears to be relatively mild. Only the simplest examples are described in this paper where it is assumed that the number of inputs is equal to the number of neurons (this restriction is for reasons of length). A review of relevant previous results on unbiased networks is included. Aside from Gaussian statistics, the main mathematical tools and ideas come from analytic geometry (analytic and subanalytic sets, the Curve Selection Lemma).