Data scale factors affect brain image model training outcomes

Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture

Artificial Intelligence

Summary

When training computer models to understand detailed images of the human brain, how the training data is spread out matters. The authors found that increasing the total number of image samples, covering more brain areas, and using bigger models all help performance. But simply using images from more people does not improve the model if the total number of images is fixed. This means that where and how many images are taken can be more important than just the number of people they come from.

What this means in practice

  • For medical image analysts: Optimize training datasets for brain image models by prioritizing sample count and spatial coverage over adding more patient subjects.
  • For machine learning engineers: Design representation learning workflows handling spatial data by decomposing data scaling into sample count, source diversity, and spatial coverage.

Authors

Christian Schiffer, Mathis Bode, Thomas Lippert, Katrin Amunts, Timo Dickscheid

Abstract

Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and a sample is an image patch at a specific spatial location. Across 93 controlled pretraining runs of a contrastive model that uses spatial proximity for supervision, we vary data allocation, compute, and model capacity over 11.6 million spatially anchored image patches from 21 human brains. Performance improves with more unique samples, broader spatial coverage, additional compute, and larger model capacity. At fixed sample count, distributing samples across one to 18 subjects produces no detectable improvement, even though representations generalize substantially better to subjects encountered during pretraining. Inter-subject variation therefore strongly affects generalization, but additional subjects provide no benefit when a fixed sample budget is distributed across more sources. These results establish sample count, source diversity, and spatial coverage as distinct axes of data scaling in spatially structured representation learning.