AI summaryⓘ
The authors study how to estimate the null space (the set of vectors mapped to zero) from a noisy matrix. They provide exact formulas and series expansions to understand errors in singular vectors and risks related to predictions from empirical data. Their approach extends to larger null spaces and they carefully analyze when their formulas converge by connecting it to special mathematical points called exceptional points. They also explore how noise and data structure affect the ranking of null directions, with theoretical proofs and Monte Carlo experiments supporting their findings. Finally, they clarify that several important quantities related to spectral analysis and noise tolerance are distinct and need separate consideration.
null spacesingular value decomposition (SVD)left singular vectorexceptional pointempirical riskgeneralization riskWishart matrixGaussian noiseMonte Carlospectral mixing
Authors
Xin Li, Jonathan Cohen, Rami Puzis
Abstract
We study null-space estimation from a noisy matrix. For a simple left null space, we first derive an exact compact expression for the error of the smallest left singular vector. We then give an all-order series for the SVD vector and projector, followed by compact and consistently truncated series forms for the fixed-realization empirical risk and conditional population generalization risk. The recursion extends to a multiple-dimensional null space by following the complete invariant subspace. The convergence radius is not inferred from an error plot: it is computed independently from the nearest complex exceptional point that joins a retained eigenvalue branch to its complement. A reduced-nullity experiment shows that moving this spectral boundary can increase the radius, although the improvement is not monotone in the retained nullity. For individually ordered null directions under Gaussian training with \(τ\geq m\), we prove that the Wishart splitting matrix \(W\) gives a strict second-order empirical ranking. Gaussian averaging equalizes the leading generalization risks at both small and very large noise, while a column-swap theorem proves strict expected generalization ranking for an isotropic signal subspace. For unequal spikes, an exact population-overlap criterion and a simultaneous \(99\%\) Monte Carlo confidence certificate explain the observed intermediate ranking. A sixth-order risk correction improves the lower-crossover estimate in the reported experiment. This equal--ranked--equal phenomenon is a finite-sample diagnostic related to spectral mixing, but its tolerance crossings, the exceptional-point radius, and the asymptotic BBP threshold are three distinct quantities.