Density-Reweighted Entropic Optimal Transport: Decoupling Geometry from Sampling Density
2026-08-17 • Machine Learning
Machine Learning
AI summaryⓘ
The authors study how to better match points between two datasets that are shaped similarly but sampled unevenly. They note that a common method called Entropic Optimal Transport (EOT) sometimes pairs points based on how densely the data is sampled rather than their actual geometric closeness. To fix this, they create a new version that can reduce the effect of sampling density, focusing more on geometry. Their method converges nicely under certain conditions and works better in simulations where sampling is uneven.
Dataset alignmentEntropic Optimal TransportSampling densityGeometric proximityTransport planRegularity conditionsPopulation-level planData correspondenceLow-dimensional structuresSimulation
Authors
Keyi Li, Yuval Kluger, Boris Landa
Abstract
Dataset alignment is a central step in data analysis across science and engineering, where the goal is to match observations between datasets. Entropic Optimal Transport (EOT) offers a computationally tractable framework for this task by encoding cross-dataset affinities in a transport plan. However, when two datasets are sampled from geometrically similar low-dimensional structures with substantially different sampling densities, the EOT plan may match points by relative sampling density rather than geometric proximity, yielding geometrically misleading correspondences. To address this issue, we propose a density-reweighted EOT framework in which the influence of sampling density on the transport plan can be discounted to a desired degree, ranging from standard EOT to alignment driven purely by underlying geometry. Under suitable regularity conditions, we establish convergence of the reweighted EOT plan to a family of population-level plans whose dependence on sampling density is made explicit. Through simulations, we show that our approach recovers geometrically faithful correspondences, improving over related EOT-based frameworks when datasets exhibit substantial sampling density disparity.