Stochastic Separability of Embedding Manifolds
2026-08-24 • Machine Learning
Machine LearningComputer Vision and Pattern Recognition
AI summaryⓘ
The authors study how objects from different categories are represented in high-dimensional spaces, like those in the brain or deep learning models. They prove mathematically that when two groups of objects have different average features and limited variation, their representations can be separated by a simple divider with high probability. This relies on a condition about how the data projects in space. Their work provides a solid theoretical foundation explaining why these object groups form low-dimensional shapes that can be separated, which helps us understand how neural networks learn to distinguish categories.
embedding manifoldhigh-dimensional spacelinear separabilitymeasure concentrationprojectionstochastic separabilityneural representationobject manifoldlaw of total expectationtwo-layer tail-bound inequalities
Authors
Liqing Zhang
Abstract
Neurobiological studies and representation learning have observed that representations of objects belonging to the same category in high-dimensional neural spaces exhibit low-dimensional object manifold characteristics, and different object manifolds are linearly separable in these neural spaces. However, these experimentally observed phenomena lack rigorous theoretical validation to date. This paper proposes a new stochastic separability theorem for embedding manifolds of two different object categories. First, we establish a projection measure concentration theorem for embedding manifolds under general conditions. We develop a new two-layer measure concentration analysis technique, which unifies two estimation bounds via the law of total expectation to derive measure concentration inequalities. Based on the measure concentration theorem, we further prove a stochastic separability theorem for embedding manifolds of two different object categories. If two datasets have distinct means and bounded total variances, their samples become linearly separable with high probability, provided that the projection direction satisfies a non-singularity condition. The main contributions of this paper are twofold: 1. We prove the projection concentration properties of embedding manifolds in high-dimensional spaces by using two-lawyer tail-bound inequalities. 2. We identify a non-singularity condition for the stochastic separability between embedding manifolds, and rigorously prove the stochastic projection separability theorem. The theorem not only uncovers geometric and statistical properties of the object embedding manifolds, but also provides a novel mechanism for representation learning in deep networks.