Back

Data-driven Sampling Strategies for Fine-Tuning Bird Detection Models

Bernard, C.; McEwen, B.; Cretois, B.; Glotin, H.; Stowell, D.; Marxer, R.

2025-10-04 bioinformatics
10.1101/2025.10.02.679964 bioRxiv
Show abstract

Passive Acoustic Monitoring has emerged as a promising tool for collecting ecological data, particularly in the context of bird population monitoring. Bird species can be automatically identified using pre-trained models, such as BirdNET. The performance of these models can be significantly improved through fine-tuning with annotated samples recorded in the specific acoustic conditions in which the microphones are deployed. However, PAM collects vast amounts of data, and annotating bird vocalizations requires specialized expetise. As a result, only a very small portion of the recordings can be effectively labeled. Selecting the most relevant samples to annotate in order to maximize performance in model fine-tuning remains a significant challenge. First, a regularization technique addresses the challenge of class imbalance during model fine-tuning. Next, a data-driven methodology is developed, introducing the influence score, which quantifies the impact of individual training samples on model performance to inform sampling strategies. A linear model is proposed to estimate the influence score for generalization to unseen data. Finally, several sampling strategies are compared, based on acoustic indices and predictions of the pre-trained model. Together, these contributions enable the identification of efficient annotation strategies to overcome the challenges of limited annotation resources in large-scale passive acoustic monitoring.

Published in The Journal of the Acoustical Society of America (predicted rank #4) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.