Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study
Machine Learning
Summary
The authors studied how to create a foundation model using very complex and mixed types of data from nuclear fusion experiments. This data comes from over 20 different sensors with very different speeds and formats, like single measurements, sound-like spectrograms, and images. They explore how to balance the details we get over time versus in frequency to best represent these mixed signals. Their work helps guide how to handle and use multi-type scientific data efficiently, which could improve control systems in nuclear fusion research.
Authors
Nathaniel Chen, Kouroche Bouchiat, Peter Steiner, Azarakhsh Jalalvand, SangKyeun Kim, Egemen Kolemen
Abstract
Training effective foundation models requires massive and organized datasets, yet scientific domains such as nuclear fusion present unique challenges due to largely heterogeneous and sparse data. Here we characterize the data used in developing such a model: with over 20 sensor types spanning 5 orders of magnitude in sampling rate, mixed tensor structures (point measurements, spectrograms, images), and nonstationary physics. We analyze our input complexity and discuss trade-offs between temporal context and frequency resolution. Our analysis provides a template for representing multi-modal fluctuation data at scale, with implications for both multi-modal control systems and nuclear fusion.