Visual and auditory deep learning models capture neural representations of naturalistic social interaction in the superior temporal sulcus
Peleg, I.; Almog, S.; Kadushin, M.; Grosbard, I.; Guy, N.; Tavor, I.; Yovel, G.
Show abstract
The superior temporal sulcus (STS) is selectively responsive to multimodal social interactions. Yet studies so far have relied on pre-defined, simplified stimuli or features to uncover the type of information that drives STS activity. We hypothesized that high-dimensional representations from visual and auditory deep learning models would better predict STS responses to naturalistic social interactions. We used self-supervised visual and auditory deep learning models to extract representations of movie frames and audio, respectively, of a TV series participants watched during fMRI scanning. Voxel-wise encoding models of a joint visual-auditory representation outperformed human-made social-affective annotations in predicting STS. Variance partition further revealed visual-auditory posterior-to-anterior gradient within the STS. To interpret what these models encode, we applied Principal Component Analysis to the encoding model weights. In both the visual and auditory models the first dimension tracked social interaction and peaked in the STS, indicating that social interaction is a dominant dimension of STS representation across both modalities. We conclude that the STS represents naturalistic social interaction in a multimodal manner, integrating visual and auditory information, and that visual and auditory deep learning models capture key representational properties of these responses.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Encoding neural representations of time-continuous stimulus-response transformations in the human brain with advanced deep neural networks 95%
- Representational similarity learning reveals a graded multi-dimensional semantic space in the human anterior temporal cortex 94%
- The Individualized Neural Tuning Model: Precise and generalizable cartography of functional architecture in individual brains 94%
Similar papers in this journal
Similar papers in this journal
- Non-linear manifold learning in fMRI uncovers a low-dimensional space of brain dynamics 94%
- Prediction of individual melodic contour processing in sensory association cortices from resting state functional connectivity 94%
- Simultaneous Modeling of Reaction Times and Brain Dynamics in a Spatial Cuing Task 93%
Similar papers in this journal
- Prediction, Syntax and Semantic Grounding in the Brain and Large Language Models 94%
- Identifying task-relevant spectral signatures of perceptual categorization in the human cortex 93%
- Tracking cortical representations of facial attractiveness using time-resolved representational similarity analysis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.