Extracting extended vocal units from two neighborhoods in the embedding plane
Lorenz, C.; Hao, X.; Tomka, T.; Ruettimann, L.; Hahnloser, R.
Show abstract
Annotating and proofreading data sets of complex natural behaviors are tedious tasks because instances of a given behavior need to be correctly segmented from background noise and must be classified with minimal false positive error rate. Low-dimensional embeddings have proven very useful for this task because they provide a visually appealing overview of a data set in which relevant clusters appear spontaneously. However, low-dimensional embeddings introduce errors because they fail to preserve high dimensional distances; and embeddings represent only objects of fixed dimensionality, which conflicts with natural objects such as vocalizations that have variable dimensions stemming from their variable durations. To mitigate these issues, we introduce a semi-supervised method for simultaneous segmentation and clustering of vocalizations. We define vocal units of a given type in terms of two density-based regions in low-dimensional embedding space, one associated with onsets and the other with offsets. We demonstrate our approach on the task of clustering adult zebra finch vocalizations embedded into the 2d plane with UMAP. We show that two-neighborhood (2N) extraction allows the identification of short and long vocal renditions from continuous data streams without initially committing to a particular segmentation of the data. Also, 2N vocal extraction achieves much lower false positive error rate than approaches based on a single defining region.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Bird song comparison using deep learning trained from avian perceptual judgments 95%
- Improving the workflow to crack Small, Unbalanced, Noisy, but Genuine (SUNG) datasets in bioacoustics: the case of bonobo calls 93%
- Capturing the songs of mice with an improved detection and classification method for ultrasonic vocalizations (BootSnap) 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.