Automatic Auditory Streaming Restores Missing Temporal Modulations in Echoic Speech
Gao, J.; Fang, M.; Chen, H.; Ding, N.
Show abstract
Human listeners can reliably recognize speech in adverse listening environments, and previous studies have identified that reliable neural encoding of slow temporal modulations in speech is essential for speech recognition. Recent behavioral studies demonstrate that long-delay echoes, which are rare in physical environments but common during online conferencing, can eliminate critical temporal modulations. These echoes, however, barely affect speech intelligibility, and here we investigate the underlying neural mechanisms. MEG experiments demonstrate that cortical activity can effectively track the temporal modulations eliminated by an echo, which cannot be explained by basic neural adaptation mechanisms such as synaptic depression, gain control, and adaptive filtering. Instead, the cortical response to echoic speech is better explained by a model that segregates speech from its echo than a model that encodes echoic speech as a whole. The speech segregation effect is observed even when attention is diverted, but disappears when speech segregation cues in the spectro-temporal fine structure are degraded. Altogether, these results strongly suggest that the auditory system can automatically segregate speech and its echo and encode them as two auditory streams, providing a potential neural basis for reliable speech recognition in echoic environments.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Spatiotemporal brain hierarchies of auditory memory recognition and predictive coding 96%
- Neural attentional-filter mechanisms of listening success in middle-aged and older individuals 96%
- Convergent neural signatures of speech prediction error are a biological marker for spoken word recognition 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Linguistic processing of task-irrelevant speech at a Cocktail Party 97%
- Distorted neural signal dynamics create hypersensitivity to background noise after hearing loss 96%
- Differential destinations, dynamics, and functions of high- and low-order features in the feedback signal during object processing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.