Humans rely more on talker identity than temporal coherence in an audiovisual selective attention task using speech-like stimuli
Cappelloni, M. S.; Mateo, V. S.; Maddox, R. K.
Show abstract
Audiovisual integration of speech can benefit the listener by not only improving comprehension of what a talker is saying but also helping a listener pick a particular talkers voice out of a mix of sounds. Binding, an early integration of auditory and visual streams that helps an observer allocate attention to a combined audiovisual object, is likely involved in audiovisual speech processing. Although temporal coherence of stimulus features across sensory modalities has been implicated as an important cue for non-speech stimuli (Maddox et al., 2015), the specific cues that drive binding in speech are not fully understood due to the challenges of studying binding in natural stimuli. Here we used speech-like artificial stimuli that allowed us to isolate three potential contributors to binding: temporal coherence (are the face and the voice changing synchronously?), articulatory correspondence (do visual faces represent the correct phones?), and talker congruence (do the face and voice come from the same person?). In a trio of experiments, we examined the relative contributions of each of these cues. Normal hearing listeners performed a dual detection task in which they were instructed to respond to events in a target auditory stream and a visual stream while ignoring events in a distractor auditory stream. We found that viewing the face of a talker who matched the attended voice (i.e., talker congruence) offered a performance benefit. Importantly, we found no effect of temporal coherence on performance in this task, a result that prompts an important recontextualization of previous findings.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Weak neural signatures of spatial selective auditory attention in hearing-impaired listeners 96%
- Magnified interaural level differences enhance spatial release from masking in bilateral cochlear implant users 96%
- Gender and speech material effects on the long-term average speech spectrum, including at extended high frequencies 96%
Similar papers in this journal
- Speech auditory-motor adaptation lacks an explicit component: reduced adaptation in adults who stutter reflects limitations in implicit sensorimotor learning. 96%
- Perceived multisensory common cause relations shape the ventriloquism effect but only marginally the trial-wise aftereffect 96%
- Top-down modulation of neural envelope tracking: the interplay with behavioral, self-reported and neural measures of listening effort 96%
Similar papers in this journal
- Visualizing sounds: training-induced plasticity with a visual-to-auditory conversion device 96%
- Auditory-cognitive determinants of speech-in-noise perception: structural equation modelling of a large sample 96%
- The role of temporal coherence and temporal stability in the build-up of auditory grouping 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.