Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD
2026-07-27 • Machine Learning
Machine Learning
AI summaryⓘ
The authors study how well we can estimate certain measures that compare different data distributions using kernels, like MMD, HSIC, and KSD. They show that the best possible accuracy for these estimates improves at a rate proportional to 1 divided by the square root of the sample size (n^{-1/2}), even in very general settings beyond simple Euclidean spaces. This means their findings confirm that the known estimation speed is actually the best we can do under fairly broad conditions. The work also applies to related tasks like estimating mean embeddings and covariance operators. Overall, the authors clarify the fundamental limits of estimating these kernel-based differences.
kernel discrepanciesmaximum mean discrepancy (MMD)Hilbert-Schmidt independence criterion (HSIC)kernel Stein discrepancy (KSD)minimax lower boundparametric convergence ratereproducing kernel Hilbert space (RKHS)mean embeddingcross-covariance operatortopological spaces
Authors
Jose Cribeiro-Ramallo, Florian Kalinke, Zoltán Szabó
Abstract
Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others. Their fastest estimators are known to converge at a parametric rate---$n^{-1/2}$---under mild conditions. While this rate is known to be minimax optimal on $\mathbb R^d$ under strict assumptions with bounded kernels, little is known about its optimality beyond the finite-dimensional Euclidean setting with unbounded kernels. In this work, we prove that the minimax lower bound of estimation of the most popular kernel discrepancies (maximum mean discrepancy, Hilbert-Schmidt independence criterion and kernel Stein discrepancy; MMD, HSIC, KSD) is $n^{-1/2}$ on general topological spaces, and under mild assumptions on the kernel; the same rates are shown (as corollaries) to hold for the estimation of the mean embedding and the centered cross-covariance operator. Our results settle the question of optimal estimation of these kernel discrepancies.