Individual Bird Identification by Modeling Temporal Structure in Bioacoustic Embeddings
Gallego, J.; Martinez-Vargas, J. D.; Lopez, J. D.
Show abstract
O_LIIdentifying individual animals from vocalizations is an emerging research area in computational bioacoustics. This non-invasive approach reduces reliance on physical capture and tagging for wildlife monitoring. Recent advances in this area leverage deep learning and bioacoustic foundation models adapted from species-level classifiers. However, these models typically rely on fixed-window inputs and may not fully capture temporal structure across extended or complex songs, which can contain information relevant to individual discrimination. C_LIO_LIHere, we evaluate whether modeling sequences of pretrained bioacoustic embeddings improves acoustic individual identification. We developed a framework that integrates transfer-learned spectrotemporal representations from BirdNET with a lightweight long short-term memory network. Unlike static baselines that use either the first embedding or an average over all embeddings, our approach processes time-ordered embedding sequences, allowing the classifier to use information distributed across multiple windows. C_LIO_LIWe evaluated the system using 87,865 vocalizations from 352 individuals across seven species. We used four publicly available vocalization datasets with individual-level labels, covering durations from 0.76 s (short calls) to 27.6 s (prolonged songs). Across five random seeds with stratified partitions, the framework achieved mean test accuracies between 93.9% and 98.3% and macro-F1 scores ranging from 93.2% to 98.3%, without data augmentation. The clearest gains over the strongest static baseline were observed for the great tit and the great spotted kiwi, reaching +1.3 and +2.8 percentage points, respectively. C_LIO_LIOur results indicate that the contribution of temporal sequence modeling depends on vocalization structure, rather than providing uniform evidence that chronological order drives performance. The benefit of recurrent aggregation was limited or absent for short calls but evident for long vocalizations with multiple informative windows. By combining bioacoustic foundation models with lightweight recurrent modeling, this approach provides a scalable, CPU-efficient tool for autonomous wildlife monitoring, particularly for species with extended or structurally complex vocalizations. C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.