A brain-inspired algorithm enhances automatic speech recognitionperformance in multi-talker scenes
Boyd, A. D.; Sen, K.
Show abstract
Modern automatic speech recognition (ASR) systems are capable of impressive performance recognizing clean speech but struggle in noisy, multi-talker environments, commonly referred to as the "cocktail party problem." In contrast, many human listeners can solve this problem, suggesting the existence of a solution in the brain. Here we present a novel approach that uses a brain inspired sound segregation algorithm (BOSSA) as a preprocessing step for a state-of-the-art ASR system (Whisper). We evaluated BOSSAs impact on ASR accuracy in a spatialized multi-talker scene with one target speaker and two competing maskers, varying the difficulty of the task by changing the target-to-masker ratio. We found that median word error rate improved by up to 54% when the target-to-masker ratio was low. Our results indicate that brain-inspired algorithms have the potential to considerably enhance ASR accuracy in challenging multi-talker scenarios without the need for retraining or fine-tuning existing state-of-the-art ASR systems.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception 94%
- Modulation masking and fine structure shape neural envelope coding to predict speech intelligibility across diverse listening conditions 94%
- Magnified interaural level differences enhance spatial release from masking in bilateral cochlear implant users 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.