Methods for comparing shapes and patterns in data using landmarks
Statistical Inference for Persistence Diagrams via Landmark Embeddings: Minimax Theory and Finite Approximation
Computational Geometry
Summary
Measuring differences in complex shapes called persistence diagrams is tricky because averaging can hide important details. The authors study ways to summarize and compare these diagrams using landmarks, which simplify the information while keeping key features. They develop mathematical tools to test differences between groups and build reliable confidence estimates. Their approach works under realistic conditions, like some parts missing or shifted, and they show how to ensure tests remain powerful even with limited data. They apply their methods to brain connectivity data to demonstrate practical use.
persistence diagramsHilbert space embeddingslandmark embeddingsstatistical inferencepopulation meantwo-sample testingconfidence setsminimax theorytemplate diagramstransport separation
Authors
Pramita Bagchi, Sushovan Majhi, Atish Mitra, Žiga Virk
Abstract
Hilbert-space embeddings enable inference for populations of persistence diagrams, but separation between individual diagrams need not survive population averaging. We develop a framework for inference on population mean embeddings, with particular attention to the additive landmark representations PLACE and PALACE. Treating each diagram as one independent observation, we apply Hilbert-space limit theory to obtain covariance estimators, two-sample tests, and confidence balls under suitable moment conditions, without requiring a lower-distortion bound. For additive embeddings, we identify the population mean as an embedding of the mean counting measure and show that geometric separation of these measures alone cannot guarantee uniform testing power. We then introduce a model with latent template diagrams, missing features, and location perturbations. Under common or feature-specific prevalence conditions, a diagram-level lower-distortion certificate yields explicit lower bounds on population mean separation. These margins provide finite-sample uniform power guarantees, and an additional information-divergence comparison gives matching sample-complexity bounds over restricted scale ranges. Confidence sets yield lower bounds on transport separation of population mean measures and exclusion guarantees for specified structured alternatives. We also quantify how orthogonal truncation changes the certified signal and the approximation allowance needed for confidence sets targeting the full embedding, relating sample size, retained coordinates, and template separation. Simulations examine calibration, power, and coverage, and an analysis of resting-state connectivity from the Autism Brain Imaging Data Exchange illustrates the procedures.