An Asymptotic Analysis of the Shapley Value for Dataset Valuation
2026-07-03 • Computer Science and Game Theory
Computer Science and Game Theory
AI summaryⓘ
The authors study how to value different datasets using the Shapley value, a method originally made to fairly divide credit among contributors. They focus on cases where the value depends smoothly on the data, using tools from math called RKHS mean embeddings. They show that although the Shapley value is usually complicated to compute, it can be closely approximated by a simpler main term when there are many datasets. This simpler term helps understand and estimate the contribution of each dataset more easily, especially when dealing with lots of data sources.
Shapley valueDataset valuationReproducing kernel Hilbert spaceRKHS mean embeddingAsymptotic analysisEmpirical distributionFirst-order approximationCombinatorial definitionData contributionEstimator benchmarking
Authors
Mélissa Tamine, Benjamin Heymann, Maxime Vono, Patrick Loiseau
Abstract
We propose an asymptotic analysis of the Shapley value in a dataset valuation setting in which utilities are modeled as smooth functionals of empirical distributions via reproducing kernel Hilbert space (RKHS) mean embeddings. We prove that, despite its combinatorial definition, the Shapley value of a data source is asymptotically captured by a simple leading term. This term can be interpreted as the first-order contribution of a dataset relative to the surrounding data population. It also identifies the scale of the Shapley value as the number of data sources grows and provides a framework for analyzing existing Shapley value estimators. Moreover, for practitioners working with large numbers of datasets, the leading term becomes a tractable reference against which Shapley value approximations can be benchmarked.