Earth observation models change more when fine-tuned than natural image models
Reuse or Relearn? A Spectral View of Earth Observation Foundation Models
Machine Learning
Summary
Fine-tuning is a common method to adapt big AI models to specific tasks. The authors studied if fine-tuning changes Earth observation (EO) AI models a lot or keeps their original knowledge. They found EO models change more extensively compared to models trained on everyday photos. This means EO models often learn new information rather than just reusing what they already knew. Their findings suggest it’s important to evaluate these models not just by how accurate they are, but also by how much of their original knowledge is preserved.
What this means in practice
- •For earth observation teams: Assess and choose EO foundation models by considering how much pretrained knowledge they preserve when fine-tuned, to optimize adaptation cost and performance.
- •For machine learning engineers: Design adaptation strategies that update fewer parameters for models with well-preserved pretrained parts, improving efficiency without losing accuracy.
Authors
Mehmet Ozgur Turkoglu, Valerio Marsocci, Dominik J. Mühlematter, Dominik Senti, Konrad Schindler, Helge Aasen
Abstract
Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy alone cannot answer: does fine-tuning reuse the pretrained representation, or does it relearn a new one? We study this with spectral diagnostics that compare a model before and after adaptation, quantifying how well its dominant singular subspaces are preserved, how broadly the weight update is distributed, and how large it is. Using natural image models such as CLIP and DINO as a reference, we find that, under the evaluated fine-tuning settings, EO models undergo far larger, higher-rank updates and retain much less of their pretrained structure, so their downstream performance is often obtained with substantial changes to the pretrained weight structure. The diagnostics further provide insight into how cheaply a model can be adapted: where the pretrained subspaces are preserved, adapting a small fraction of the parameters can match full fine-tuning, and where they are not, it can fall behind. More broadly, foundation models, and EO foundation models in particular, should be assessed not only by benchmark accuracy, but also by how reusable their pretrained representation is.