The Psychoacoustics of Automatic Speech Recognition
Weerts, L.; Clopath, C.; Goodman, D. F. M.
Show abstract
Deep neural networks have had considerable success in neuroscience as models of the visual system, and recent work has suggested this may also extend to the auditory system. We tested the behaviour of a range of state of the art deep learning-based automatic speech recognition systems on a wide collection of manipulated sounds used in standard human psychometric experiments. While some systems showed qualitative agreement with humans in certain tests, in others all tested systems diverged markedly from humans. In particular, all systems used spectral invariance, temporal fine structure and speech periodicity differently from humans. We conclude that despite some promising results, none of the tested automatic speech recognition systems can yet act as a strong proxy for human speech recognition. However, we note that the more recent systems with better performance also tend to better match human results, suggesting that continued cross-fertilisation of ideas between human and automatic speech recognition may be fruitful. Our open source toolbox allows researchers to assess future automatic speech recognition systems or add additional psychoacoustic measures.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception 97%
- Representations of fricatives in sub-cortical model responses: comparisons with human consonant perception 96%
- Gender and speech material effects on the long-term average speech spectrum, including at extended high frequencies 96%
Similar papers in this journal
- Speech-driven Facial Animations Improve Speech-in-Noise Comprehension of Humans 97%
- The Effect on Speech-in-Noise Perception of Real Faces and Synthetic Faces Generated with either Deep Neural Networks or the Facial Action Coding System 96%
- Electromyographic Correlates of Effortful Listening in the Vestigial Auriculomotor System 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.