Skip to content

OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations

Conference: NeurIPS 2026 Datasets & Evaluations Track Spotlight (not the ordinary main track)
arXiv: 2609.34839
Area: Audio & Speech
Keywords: dolphin whistles, bioacoustics, longitudinal dataset, self-supervised learning, linear probing

TL;DR

OpenWhistle organizes years of underwater recordings from one dolphin group into a pretraining corpus containing approximately 180,000 whistles and an expert-annotated subset of 8,354 whistles, establishing session-split classification and detection benchmarks with frozen-representation linear probes; corpus-pretrained Wav2Vec2.0 achieves 81.1% classification accuracy and 75.6% detection mAP, without decoding the meaning of dolphin communication.

Background & Motivation

Large-scale bioacoustic resources have made animal sound detection and species classification systematically comparable, but resources such as Xeno-Canto and BirdSet primarily provide broad coverage across species. Many short recordings of a species are not equivalent to years of continuous interactions among the same individuals: studying vocal identity, whistle exchanges, social relationships, and acoustic drift requires knowledge of the animals living together and the temporal relationships between successive calls. Dolphin signature whistles (SW) are particularly relevant because their frequency contours are associated with individual identity and dolphins exhibit vocal learning and imitation; existing datasets, however, are often small, access-restricted, or limited to isolated clips and coarse labels.

OpenWhistle takes advantage of the semi-natural observation conditions at Dolphin Reef: fixed hydrophones can record a known group over time while dolphins retain access to the open sea, preserving environmental sounds, overlapping vocalizations, and variable signal-to-noise ratios. This is not a new universal corpus covering dolphins worldwide, but a deep resource centered on one site and a group of known individuals; large unlabeled sequences support representation learning, while a smaller expert-labeled subset supports reproducible evaluation. This separation makes it possible to examine a specific question: whether learning dolphin acoustic structure directly is more useful for fine-grained within-species discrimination than transferring general audio or broad-coverage animal sound models.

The main contribution is therefore not a new network architecture, and whistle categories are not semantic labels. Core idea: connect longitudinal raw recordings, scalable whistle screening, and expert contour labels into an unlabeled pretraining corpus plus a controlled downstream benchmark, testing whether in-domain self-supervised representations distinguish subtle whistle-type differences.

Method

Overall Architecture

The input consists of underwater recordings collected at Dolphin Reef between 12 November 2019 and 28 March 2024, comprising 7,495 recording sessions and 6,271.78 hours of usable audio. Three hydrophones recorded at an original sampling rate of 96 kHz; the group included four resident Tursiops truncatus ponticus dolphins and Yosefa, an intermittently visiting female Tursiops aduncus. The paper's recurring description of five known individuals should therefore not be interpreted as five animals of one species continuously present together.

Dataset construction first performs binary whistle presence detection and then branches according to the intended use. One branch merges nearby positive segments into sequences retaining temporal context, producing approximately 114 hours of pretraining audio; the other estimates fundamental-frequency contours, assigns initial categories with ARTwarp, and applies expert inspection to produce 8,354 annotated whistles. The branches have different quality requirements: pretraining preserves real noise and overlap, whereas expert annotation favors clear vocalizations with little overlap.

Evaluation compares logistic regression linear probes on frozen audio representations rather than training an end-to-end detection network. Classification receives an already isolated whistle and predicts one type; detection receives a fixed 0.5-second segment and predicts the presence of each type through multi-label outputs. Wav2Vec2.0 is first self-supervised pretrained on OpenWhistle and subsequently evaluated under the same downstream probing protocol as handcrafted features, AVES, and BioLingual. The contribution is primarily the dataset and evaluation protocol; the following explanation follows the actual processing steps without presenting a dataset directory or task checklist as a network architecture diagram.

Key Designs

1. Whistle presence detection: screen usable vocalization segments from longitudinal recordings

Annotating every whistle in thousands of recording hours would be expensive, so the authors first solve the simpler binary question of whether a whistle is present. Each non-overlapping 0.4-second audio window is converted to a log-power spectrogram using a 1,024-sample Blackman window, a 1,024-point FFT, and a hop of 512 samples, then cropped to 2โ€“22 kHz, minโ€“max normalized, and resized to 224ร—224 pixels. The input is replicated across three channels and normalized with ImageNet statistics to reuse ImageNet-pretrained VGG16. The original classifier is replaced by a lightweight fully connected head with hidden dimensions of 50 and 20, and the entire network is fine-tuned rather than only the final classifier.

This detector only decides whether at least one whistle occurs in a window; it does not identify whistle type, the vocalizing individual, or meaning. Spectrogram whistle structure reduces the amount of subsequent processing, but screening also introduces a selection boundary: if a low-SNR whistle is missed, later sequence construction and expert annotation do not automatically recover it. Consequently, even 97.99% recall cannot establish that all animal vocalizations were recorded in the corpus; the authors estimate that approximately 3,700 whistles were not captured.

2. Longitudinal sequence construction: retain temporal relationships instead of exporting only isolated samples

Adjacent positive detection windows are first concatenated into continuous segments, and segments separated by less than 6 seconds are merged into the same sequence. The retained material includes not only whistles but also surrounding acoustic context and neighboring vocalizations, allowing unlabeled pretraining to encounter continuous acoustic variation during interactions. The reported corpus contains 33,267 sequences, approximately 114 hours of audio, and approximately 180,000 whistles; the whistle count is estimated from duration and mean whistle length rather than obtained through exhaustive manual counting.

The main text reports a mean sequence duration of 12.95 seconds, a standard deviation of 19.9 seconds, a range of 5โ€“246 seconds, and a mean segment interval of 4.11 seconds. Appendix B, however, states that sequences of 2โ€“20 seconds are retained for sequence-level summaries without explaining the relationship between these sets; the final pretraining samples should not therefore be uniformly described as 2โ€“20 seconds long. Longitudinal coverage also does not imply uniform, uninterrupted observation over five years: recordings concentrate in particular periods, 2022 is entirely missing, and both offshore foraging and human activity influence which sounds enter the dataset.

3. Expert contour annotation: combine scalable categorization with human acoustic judgment

Whistle-type discrimination depends on subtle frequency contours rather than merely detecting a bright line in a spectrogram. The authors estimate fundamental frequency (F0) with a dolphin-specific CREPE model, applying frequency compression with compress=20 and multiplying decoded F0 values back by the same factor to accommodate whistles above the original CREPE pitch range. F0 is estimated every 5 ms with the weighted_argmax decoder; contours are flagged as low-confidence when fewer than 5% of frames exceed a confidence threshold of 0.05. The paper specifies flagging, not that every low-confidence contour is discarded.

ARTwarp then combines dynamic time warping (DTW) with an adaptive resonance theory network to establish initial assignments from contour similarity and manually annotated templates; vigilance is set to 90 and other parameters remain at their defaults. DTW helps compare contours with similar shapes but different speeds or lengths, while experts inspect spectrograms to correct assignments and ambiguity rather than treating clustering outputs as ground truth. The resulting 10 categories comprise 7 signature-whistle types with 7,624 whistles, or 91.3%, and 3 non-signature-whistle (NSW) types with 730 whistles, or 8.7%. Seven signature-whistle types do not imply seven individuals present during recording: the appendix also notes signature whistles associated with individuals no longer present; together with imitation, this prevents type prediction from directly substituting for current caller identification.

4. Session-level linear probing: separate representation quality from downstream model capacity

The classification set contains 3,000 balanced instances drawn from the 6 best-represented annotated classes; rare classes are excluded because too few examples are available. Detection divides continuous recordings into 0.5-second windows, covering 7 classes with 400 instances per class and adding 2,800 background segments. Detection uses an independent binary label for each type, permitting multiple types in a segment; background is the all-zero vector, not an eighth vocalization category, and the paper does not detail deduplication when sampling multi-label instances. The 6 classification classes, 7 detection classes, and 10 fully annotated classes are distinct scopes and should not be interchanged.

All downstream data are assigned by recording session to training, validation, and test splits in a 70% / 15% / 15% ratio, while approximately preserving class distributions. Sounds from one session cannot enter both probe training and testing, reducing shortcuts from equipment, background conditions, and contiguous acoustic context. This rule explicitly constrains downstream splits only; the paper does not state whether self-supervised pretraining excludes downstream validation and test sessions or their corresponding audio. Downstream session disjointness therefore does not establish that pretraining never encountered the test source; the in-domain advantage should retain this evaluation boundary rather than being presented as fully leakage-free inductive generalization across sessions.

Loss & Training

The VGG16 corpus-screening model uses cross-entropy, Adam, a learning rate of \(10^{-5}\), a batch size of 4, early stopping, and ReduceLROnPlateau. The appendix reports 53,828 training, 5,980 validation, and 16,708 test windows; the first two sum to 59,808, corresponding to the fine-tuning dataset size stated in the main text. The binary screening supervision, subsequent whistle-type detection benchmark, and self-supervised pretraining objective operate at different levels.

Wav2Vec2.0 uses a standard self-supervised setup, learning representations by predicting quantized latent targets from masked contexts rather than using expert whistle types as pretraining supervision. The appendix specifies two codebooks of size 320 and a Gumbel-softmax temperature that decays exponentially from 2.0 to 0.5; the architecture follows Wav2Vec2.0 base apart from the feature encoder adaptation for a higher sampling rate. Model audio is configured at 44.1 kHz, with the encoder adapted relative to the conventional 16 kHz speech setting to preserve relative temporal resolution; the paper does not list the exact convolutional strides or the resampling details connecting 96 kHz acquisition to this input. Because the paper does not fully specify its pretraining loss equation, a generic Wav2Vec2.0 equation is not supplied here as the paper's exact formulation.

Pretraining runs for 400k steps on 32 V100 GPUs, with a per-device batch size of 4 and 2 gradient accumulation steps, giving an effective batch size of 256 audio segments. AdamW uses \(\beta_1=0.9\), \(\beta_2=0.98\), \(\epsilon=10^{-6}\), a learning rate of \(5\times10^{-4}\), linear decay, 32k warmup steps, weight decay of 0.01, and mixed precision. This configuration demonstrates that the corpus supports self-supervised training, but it is not a low-compute setup, and the external models are not compared under matched training budgets.

Downstream logistic regression uses lbfgs with a maximum of 20,000 iterations; inverse regularization strength \(C\in\{0.1,1.0,10.0\}\) is selected on validation data. Classification reports accuracy, while detection reports mAP, the mean of average precision across whistle types; table values use a percentage scale throughout. The test set is resampled with replacement 1,000 times to report means and standard deviations; these error bars are neither variation across pretraining seeds nor 95% confidence intervals.

Key Experimental Results

Main Results

Table 1 corresponds to source Table 2: results under the same frozen-representation linear probing protocol, reported as mean ยฑ bootstrap standard deviation in percentage units.

Representation / model Pretraining source Classification accuracy Detection mAP
Chance level None 16.7 8.3
Spectral features None 34.9 ยฑ 2.2 26.3 ยฑ 1.1
MFCCs None 45.6 ยฑ 2.4 33.6 ยฑ 1.8
Mean spectrogram None 55.6 ยฑ 2.4 47.7 ยฑ 2.1
AVES-core AudioSet, FSD50K 68.0 ยฑ 2.2 57.4 ยฑ 2.1
BioLingual AnimalSpeak audioโ€“text 71.3 ยฑ 2.1 66.5 ยฑ 2.2
AVES-bio Labeled as animal vocalizations in the source: AudioSet, VGGSound 75.1 ยฑ 2.1 65.0 ยฑ 2.3
Wav2Vec2.0 OpenWhistle 81.1 ยฑ 1.8 75.6 ยฑ 2.0

Relative to Mean spectrogram, the strongest handcrafted baseline, Wav2Vec2.0 gains 25.5 percentage points in classification and 27.9 percentage points in detection. Selecting the strongest off-the-shelf model separately for each task, the gains are 6.0 percentage points over AVES-bio for classification and 9.1 percentage points over BioLingual for detection. The source narrative describes AVES-bio as the strongest off-the-shelf model on both tasks, but Table 2 gives BioLingual 66.5 mAP versus AVES-bio's 65.0; the table values are retained here and the discrepancy is explicit. The detection Chance level of 8.3 is reproduced as reported; the paper does not explain its calculation, so it is not revised based on the number of classes.

Ablation Study

The paper provides no module-removal, dataset-size, or sampling-rate ablations; Table 2 summarizes construction quality and error analyses from the main text and Appendices B and D, not controlled ablation results.

Analysis target Source evidence Implication for interpreting results
Binary corpus-screening detector 16,708 test windows; precision 96.52%, recall 97.99% Reliable screening does not ensure full coverage of low-SNR whistles
WMMSD binary clip check No retraining; F1 = 0.904 External check of the binary screening model, not the type classifier
WMMSD delphinid / clear-noise proxy check No retraining; F1 = 0.907 Proxy task differs from the official seven-class detection benchmark
DCLDE proxy subset check No retraining; F1 = 0.966 Insufficient to establish cross-site transfer of type representations
Expert subset quality 8,354 whistles; mean SNR 13.24 dB; mean duration 0.84 seconds Favors clear samples and does not represent the full noise distribution
Type confusion Appendix D identifies more frequent confusion between Nana and Yosefa SW; no readable numerical rate provided Similar contours and within-class variability remain difficult

Key Findings

  • General bioacoustic representations clearly outperform handcrafted features, and in-domain pretraining improves further under this protocol; architecture, pretraining source, and compute are not separately controlled.
  • AVES-bio is stronger than BioLingual in classification but weaker in detection; no off-the-shelf model dominates both tasks.
  • High classification accuracy neither establishes correct caller identification for every vocalization nor reveals whistle meaning; the evaluated labels are acoustic categories.
  • External F1 scores for binary screening and downstream seven-class detection mAP measure different tasks and cannot be combined into one cross-domain performance claim.

Highlights & Insights

  • The resource prioritizes within-species depth rather than the number of species covered. Sequence context and known group histories enable future studies of temporal change, beyond classification of isolated sounds.
  • Automated screening, contour categorization, and expert correction concentrate human judgment where fine-grained knowledge is most needed. The approach can transfer to other high-frequency animal sounds, but detector omissions must be included in the description of dataset bias.
  • Frozen representations with simple linear probes reduce confounding from downstream head capacity. They reveal whether representations are directly usable, but do not replace end-to-end fine-tuning or strictly unseen-source evaluation.

Limitations & Future Work

  • Data come from one site and a small group, including an intermittent visitor, with uneven temporal coverage and no recordings in 2022. Generalization across sites, groups, and years requires independent validation.
  • Expert labels cover approximately 5 months and are highly imbalanced, with some types having fewer than 100 examples. Balanced six-class classification does not evaluate the natural distribution of all ten categories.
  • The expert set favors clear sounds with little overlap, while pretraining retains complex environments. Stratified evaluation of low-SNR, overlapping, rare-type, and long continuous-stream inputs is needed to connect controlled scores to practical monitoring utility.
  • Downstream session separation is valuable, but pretraining overlap with test sources is unspecified. A setup strictly excluding test sessions from pretraining should be reported separately from one allowing unlabeled target-domain audio.
  • The main-text sequence range of 5โ€“246 seconds and the appendix's 2โ€“20-second retention rule require clarification; processing details for the 44.1 kHz model input are also incomplete.
  • The authors propose extending longitudinal expert labels and integrating video and behavioral context; before such supervision and dedicated evaluations exist, this benchmark should not be described as dolphin language decoding.
  • vs BirdSet / BEANS: These resources emphasize broad-coverage classification or standardized animal sound tasks, whereas OpenWhistle emphasizes longitudinal within-species structure among known individuals. They are complementary; the paper does not establish that small-group in-domain pretraining generally outperforms every broad-coverage model.
  • vs DOLPHINFREE / DCLDE: Existing resources provide smaller whistle or contour datasets, while OpenWhistle adds raw-audio scale, sequence context, and open evaluation. Its external detection checks remain proxy tasks, not comprehensive comparisons in a shared label space.
  • vs AVES / BioLingual: General animal sound pretraining supplies useful transferable representations, while OpenWhistle-trained Wav2Vec2.0 performs better on this benchmark. Attributing the advantage specifically to in-domain data requires matched architecture and compute with only the pretraining corpus changed.
  • Research leads: Year-based holdouts with the corresponding sessions strictly excluded from pretraining could test acoustic drift rather than repeated environmental context; active learning on rare types and low-SNR segments could reduce expert annotation needs. These are proposed follow-up experiments, not results established by the paper.

Rating

  • Novelty: 4/5 โ€” The combination of open, longitudinal, single-group depth is valuable; the modeling largely reuses existing methods.
  • Experimental Thoroughness: 4/5 โ€” Unified linear probes, multiple baselines, and external screening checks are provided, but strict pretraining-source separation and controlled ablations are missing.
  • Writing Quality: 4/5 โ€” Task boundaries are clear and implementation details appear in the appendix, with remaining inconsistencies in the strongest-baseline narrative and sequence-length scope.
  • Value: 4/5 โ€” A practical foundation for dolphin acoustic representation learning, with semantic understanding and cross-group transfer still requiring evidence.