Back

Talking avatars can differentially modulate cortical speech tracking in the high and in the low delta band

Riegel, J.; Schüller, A.; Jehn, C.; Wissmann, A.; Zeiler, S.; Kolossa, D.; Reichenbach, T.

2026-01-08 neuroscience
10.64898/2026.01.07.695461 bioRxiv
Show abstract

In noisy listening environments, visual cues from a speakers face can significantly boost speech compre-hension. The underlying audiovisual integration in the brain involves neural tracking of audiovisual speech features. Moreover, lip reading in silence is associated with tracking of the speech envelope in the low-delta frequency band (0.5 - 1 Hz). Recently, digital avatars have emerged that can support speech comprehen-sion. Yet, it remains unclear how the human brain integrates such artificial visual signals with natural speech. Here, we employed magnetoencephalography (MEG) to measure the neural response to a natural video, an avatar generated by deep neural networks, and a degraded video serving as a control. We demonstrate that the avatar can enhance speech-in-noise comprehension to a similar degree as the degraded video, although less than the natural video. We further identify a late response at 600 ms in the neural tracking of the audi-tory cortex in the high delta band (1 - 4 Hz) that predicts audiovisual speech comprehension. In contrast, we found that neural tracking in the low delta band is related to silent lip-reading performance. Importantly, the tracking in the low delta band evoked by the avatars is much weaker and occurs earlier than that elicited by the other audiovisual stimuli. Neural tracking in the theta band (4 - 8 Hz) is not involved in audiovisual integration. Our results show that the low delta band and the high delta band play clearly distinct roles in visual-only and audiovisual speech processing, and suggest potential avenues for further boosting the abilities of avatars to support speech comprehension. Significance StatementUnderstanding a conversational partner is essential for everyday communication. Yet, many people -- due to aging or other factors -- struggle to follow speech in noisy environments. Seeing the speakers face can greatly enhance speech comprehension, but visual cues are often unavailable, such as during public announcements or telephone conversations. Digital avatars offer a promising alternative, but how the brain integrates audiovisual information from such artificial sources remains unclear. Using magnetoencephalog-raphy (MEG), we investigated how the brain processes and integrates speech when visual information is provided by either natural or artificial (avatar-based) signals. Our findings reveal both shared and distinct neural mechanisms of audiovisual integration, providing critical insight into how visual input can support speech understanding in challenging listening conditions.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.