Active learning improves labeling for ecological sound monitoring

BioDCASE: Active Learning for Bioacoustics

Machine LearningSound

Summary

Machine learning models for monitoring animal sounds need lots of labeled examples, but labeling is costly because many recordings exist. The authors organized a challenge to compare different ways of picking which sounds to label first. They found that the best methods improved learning by over 25% compared to picking examples randomly. Combining diversity and uncertainty in choosing sounds to label worked better than using just one approach. These findings help make better use of labeling efforts for ecological sound data.

What this means in practice

  • For ecological monitoring teams: Prioritize which animal sound recordings to label to improve machine learning models efficiently and reduce labeling costs.
  • For marine biology data analysts: Use combined diversity and uncertainty criteria to select marine acoustic samples that maximize model learning gains during annotation.

Authors

Ben McEwen, Rupa Kurinchi-Vendhan, Shiqi Zhang, Lukas Rauch, Marek Herde, Sara Beery

Abstract

Ecological monitoring increasingly relies on machine learning models, whose performance depends on the quality and quantity of labelled data. However, obtaining these labels is costly, particularly in passive acoustic monitoring, where vast amounts of data are collected but only a small proportion can feasibly be annotated. Active learning addresses this bottleneck by prioritizing which samples should be labelled. However, progress is difficult to measure, because published methods are evaluated under different models, budgets, evaluation metrics and datasets. To address this challenge, we present the 2026 Active Learning for Bioacoustics BioDCASE challenge: a systematic evaluation of sampling methods designed to identify effective AL strategies. Participant methods were evaluated across four subsets composed of terrestrial and marine data. Across ten proposed sampling methods from seven teams, the top-ranked method achieved an area under the learning curve 26.4 % higher than random sampling at the same annotation budget, averaged over four data subsets. Significant variation in performance was observed across subsets, with the top-performing submission achieving a 67.1 % gain for the HSN subset over random sampling and a gain of 8 % for the ATBFL subset. Top-ranking submissions combined multiple acquisition signals, and diversity-based selection outperformed pure uncertainty sampling. Furthermore, there is evidence that transitioning from diversity-based to uncertainty-based selection and explicitly reducing redundancy within acquisition batches improve model training. There is also initial evidence that larger acquisition batch sizes may be increasingly beneficial later in the labelling process.