Back

Modelling temporal shift-invariance in self-supervised generative models improves accuracy and interpretability of species detection in soundscape recordings

Gibb, K. A.; Eldridge, A.; Shuaibu, A. L.; Simpson, I. J. A.

2025-12-11 ecology
10.64898/2025.12.09.693207 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWRealising the potential for acoustic monitoring to deliver biodiversity insight at scale requires new approaches to the automated analysis of PAM recordings that are trustworthy as well as cost-effective. Discriminative models trained on annotated species data are gaining popularity but are labour intensive, notoriously opaque and biased. Self-supervised generative models such as Variational Autoencoders (VAE) offer great potential for learning compact yet expressive representations of data, which can be used for subsequent discriminative tasks and are intrinsically interpretable. However, the default learning algorithm results in weakly discriminative data representations due to under-specification of the generative task. We propose and evaluate a novel modification to the VAE learning algorithm that models intra-frame shift-invariance. We demonstrate that this modification provides representations that are more interpretable, consistent and improve classification performance. Performance accuracy is evaluated on species detection tasks on two weakly annotated data sets across temperate and tropical terrestrial habitats and compared to leading discriminative models BirdNet and Perch, as well as the classic VAE. Whilst demonstrated in terrestrial recordings, the approach is transferable to marine, freshwater, and soil habitats. These innovations set the path for trustworthy, data and time-efficient tools to support solid ecological inference from large-scale passive acoustic monitoring surveys.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.