Vib2Sound: Separation of Multimodal Sound Sources
Akahoshi, M.; Wang, Y.; Cheng, L.; Zai, A. T.; Hahnloser, R. R. H.
Show abstract
Understanding animal social behaviors, including vocal communication, requires longitudinal observation of interacting individuals. However, isolating individual-level vocalizations in complex environments is challenging due to background noise and frequent overlaps of coincident signals from multiple vocalizers. A promising solution lies in multimodal recordings that combine traditional microphones with animal-borne sensors, such as accelerometers and directional microphones. These sensors, however, are constrained by strict limits on weight, size, and power consumption and often lead to noisy or unstable signals. In this work, we introduce a neural network-based system for sound source separation which leverages multi-channel microphone recordings and body-mounted accelerometer signals. Using a dataset of zebra finches recorded in a social setting, we demonstrate that contact sensing largely outperforms conventional microphone-array recordings. By enabling the separation of overlapping vocalizations, our approach offers a valuable tool for studying animal communication in complex naturalistic environments.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Capturing the songs of mice with an improved detection and classification method for ultrasonic vocalizations (BootSnap) 95%
- Improving the workflow to crack Small, Unbalanced, Noisy, but Genuine (SUNG) datasets in bioacoustics: the case of bonobo calls 94%
- Bird song comparison using deep learning trained from avian perceptual judgments 93%
Similar papers in this journal
- BioCPPNet: Automatic Bioacoustic Source Separation with Deep Neural Networks 97%
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 96%
- Comparison of Two-Talker Attention Decoding from EEG with Nonlinear Neural Networks and Linear Methods 93%
Similar papers in this journal
- Linear versus deep learning methods for noisy speech separation for EEG-informed attention decoding 96%
- Real-time control of a hearing instrument with EEG-based attention decoding 93%
- Speech decoding from a small set of spatially segregated minimally invasive intracranial EEG electrodes with a compact and interpretable neural network 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.