Back

The reliability of acoustic classification models for determining avian vocalisation patterns

Metcalf, O. C.; Alencar Nunes, C.; Hopping, W. A.; Lees, A. C.; Lostanlen, V.; Barlow, J.

2025-12-16 ecology
10.64898/2025.12.15.694294 bioRxiv
Show abstract

Automated detection and classification of species vocalisations offers the potential to utilise acoustic datasets across unprecedented spatial and temporal scales. However, classification algorithms inevitably generate errors, and error rates vary with context. While methods for quantifying error rates in ecoacoustics are well established, there is limited research on what level of model performance is sufficient to reliably determine avian vocalisation patterns. Using an extensive fully expert-labelled acoustic dataset from Peru (18 hours, 6 sites), we examined changes in the probability of detecting target bird vocalisations in the first hour after dawn to address three key questions: (1) How sensitive are models predicting detection probability over time to reductions in classification accuracy? (2) To what extent does aggregating detections over longer time periods impact classification accuracy? Additionally, we used the labelled dataset to assess how the creation and composition of a test dataset for assessing classifier performance can impact the reliability of accuracy metrics: (3) Are estimates of classification precision robust when test and deployment datasets are not independent and identically distributed? Our results indicate that poor classification performance--especially low precision--can lead to misleading inferences about temporal patterns of vocalisations. Aggregating classifier predictions over longer time periods improved recall but often resulted in misleading patterns of vocal behaviour by reducing precision and temporal resolution. We also demonstrate that precision can be substantially overestimated when species presences are rarer in the deployment dataset than in test data. These findings highlight the importance of cautious application of automated classification in acoustic ecology and the need for accuracy assessment methods tailored to the intended ecological analysis.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.