Largest dolphin whistle dataset aids automated communication study

OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations

Computation and Language

Summary

Studying how dolphins communicate has been hard because there were not many recordings available. The authors collected a huge amount of dolphin whistle sounds over five years from the same group of dolphins and made this data public. They also created tools to detect, break down, and classify these whistles. By training a sound-recognition model on their dataset, they showed it can understand dolphin whistles better than existing tools. This work opens up new ways to study dolphin communication in detail using machine learning.

What this means in practice

  • For marine biologists: Use the large, annotated dolphin whistle dataset and tools to analyze dolphin communication patterns and social behaviors more effectively.
  • For machine learning engineers: Train or fine-tune models on the open dolphin dataset to build specialized bioacoustic classifiers for underwater sound detection systems.

Authors

Faadil Mustun, Chiara Semenzin, Roberto Dessi, Pablo Robin Guerrero, Pierre Orhan, Alexis Emanuelli, Emanuele Rossi, Yair Lakretz, Gonzalo de Polavieja, German Sumbre

Abstract

Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (Tursiops truncatus) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.