Machine learning reveals matter density from galaxy motions and positions

Inductive Biases in Field-Level Cosmological Inference from Galaxy Catalogs

Machine Learning

Summary

Measuring how much matter is in the universe is a big challenge. This paper shows how different machine learning methods can estimate the universe’s matter density by looking at where galaxies are and how they move. Models that focus on galaxy velocities do pretty well, especially when using methods that consider relationships between galaxies. However, using galaxy data to predict other cosmic properties remains difficult. The authors also highlight that real observations need more testing because simulated galaxy motions are exact, unlike noisy data from telescopes.

What this means in practice

  • For cosmology analysis teams: Use graph neural networks to estimate matter density from galaxy positions and velocities in simulation-based cosmological studies.
  • For machine learning practitioners: Develop models that incorporate spatial relations explicitly for parameter inference on structured scientific data like galaxy catalogs.

Authors

James O. Baldwin, Shy Genel, Francisco Villaescusa-Navarro

Abstract

We perform field-level likelihood-free inference of the matter density parameter $Ω_m$ from simulated galaxy catalogs using machine learning models with differing inductive biases. Using hydrodynamic simulations from CAMELS, we examine how observable choice and architecture govern cosmological information extraction. We consider galaxy positions and line-of-sight peculiar velocities, separately and jointly, and compare permutation-invariant Deep Sets, implemented with either multilayer perceptrons (MLPs) or Kolmogorov-Arnold Networks (KANs), to graph neural networks (GNNs), which explicitly encode spatial relations. We test in-distribution and out-of-distribution (OOD) performance across simulations with different subgrid galaxy-formation prescriptions. Deep Sets infer $Ω_m$ from velocities alone with mean relative errors of approximately $18\%$ in-distribution and $\sim25\%$ OOD, with KANs and MLPs achieving comparable performance. In contrast, the same set-based approach does not yield useful $σ_8$ predictions in either in-distribution or cross-suite tests. Adding positions does not improve Deep Sets, while GNNs infer $Ω_m$ with mean relative errors of about $10\%$ in-distribution and $10$--$17\%$ OOD. These results indicate that peculiar velocities provide the dominant source of $Ω_m$ information for set-based models in this setting, while spatial information is most effectively used by architectures that explicitly encode galaxy-galaxy relations. Because the velocity inputs are exact simulated peculiar velocities, applications to survey data will require validation under realistic velocity-measurement noise, selection effects, and survey geometry.