Transformer model improves crop type mapping from satellite images

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

Computer Vision and Pattern RecognitionMachine Learning

Summary

Knowing exactly which crops grow where helps farmers and food planners. The authors created a new computer model called PAtteRNS that looks at satellite images in three ways—time, color, and space—to better identify different crop types. Their method is faster and more accurate than others, especially at detecting crop field edges. They also found that how datasets are prepared affects results a lot and suggest better standards are needed for future improvements.

What this means in practice

  • For agriculture monitoring teams: Map and identify crop types more accurately and efficiently using satellite time series data with separate attention to time, spectrum, and space.
  • For remote sensing software developers: Build advanced models for crop field edge detection that improve the quality of parcel boundary delineation in satellite image segmentation.

Authors

Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini

Abstract

The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.