Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection
Gu, N.; Lee, K.; Basha, M.; Ram, S. K.; You, G.; Hahnloser, R.
Show abstract
This paper introduces WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for human and animal Voice Activity Detection (VAD). Contrary to traditional methods that detect human voice or animal vocalizations from a short audio frame and rely on careful threshold selection, WhisperSeg processes entire spectrograms of long audio and generates plain text representations of onset, offset, and type of voice activity. Processing a longer audio context with a larger network greatly improves detection accuracy from few labeled examples. We further demonstrate a positive transfer of detection performance to new animal species, making our approach viable in the data-scarce multi-species setting.1
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- BioCPPNet: Automatic Bioacoustic Source Separation with Deep Neural Networks 96%
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 96%
- Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.