Scaling analysis evaluated for language and vision models

ScAn-Bench: Evaluating Scaling Analysis Methodology

Machine Learning

Summary

Machine learning improves when models, data, and settings are scaled up properly, but there was no clear way to check if the methods to find these scaling rules actually work. The authors created large collections of example models for language and vision-language tasks to test how scaling rules are derived and predicted. They found ways to better understand how to gather and use data for predicting model performance as they grow. This work helps future researchers and builders know which analysis methods to trust when scaling up AI models.

What this means in practice

  • For machine learning engineers: Choose more reliable methods to predict how increasing model size or data improves language and vision-language AI performance.
  • For ai infrastructure teams: Improve decisions around resource allocation for training large language and vision-language models by understanding scaling analysis methodology.

Authors

Artin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog, Neeratyoy Mallik, Jenia Jitsev, Danny Stoll

Abstract

Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 4524 and 8024 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.