Two timescale fine-tuning helps learn new features with little data

Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks

Machine Learning

Summary

Fine-tuning a pre-trained neural network on a new task with little data is common but not well understood theoretically. The authors study a simplified setting and show that tuning different parts of the network at different speeds can learn new features specific to the new task while keeping old ones intact. This approach needs fewer data samples than starting from scratch. Their work explains why fine-tuning pre-trained models is more effective than random initialization for learning from scarce data.

What this means in practice

  • For machine learning engineers: Design fine-tuning procedures that update output weights faster than hidden weights to efficiently learn new features from limited data.
  • For ai product developers: Improve adaptation of pre-trained neural networks in specialized domains with few samples by adopting the two-timescale fine-tuning method.

A theory result. No direct application yet.

Authors

Etienne Boursier, Nicolas Flammarion

Abstract

Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has $m+1$ features, $m$ of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only $\mathcal{O}(d)$ fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.