Back

Face-selective responses correlate with deep networks that learn from environment feedback

Zhou, M.; Schwartz, E.; Alreja, A.; Richardson, M.; Ghuman, A.; Anzellotti, S.

2026-02-27 neuroscience
10.64898/2026.02.25.703652 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWDeep neural networks have shown high accuracy in modeling neural responses in the visual system, but most models rely on supervised learning, which requires training on ground-truth labels that are typically unavailable in real-world settings. While unsupervised models can address this limitation, they miss another key aspect: visual representations are shaped by feedback from the environment. We introduce a reinforcement learning (RL) model of face perception that incorporates both input stimuli and feedback from the environment. Inspired by human interactions, we train the model to approach faces yielding positive interactions and avoid faces yielding negative interactions. Using intracortical electroencephalography (iEEG) data and Representational Dissimilarity Matrices (RDMs), we evaluate the models ability to account for neural responses. Our RL model performs at the same level as supervised and unsupervised models, capturing neural responses to complex visual stimuli. The findings suggest that RL models are a promising approach for understanding perception. Significance StatementUnderstanding how the brain encodes faces is central to vision science. Existing models rely on supervised learning, which requires ground-truth labels that are often unavailable in real-world settings, or on unsupervised learning, which ignores the role of environmental-feedback in shaping visual representations. We introduce a reinforcement learning (RL) model that learns through environmental feedback, simulating human interactions by associating approaching faces with positive interactions and avoiding faces with negative interactions. Using intracortical electroencephalography (iEEG) data from face-selective regions, we show that an RL model with a variational DenseNet encoder accounts for neural representations comparably to supervised and unsupervised models. Task and architecture jointly shaped representational geometry, highlighting the importance of both learning objective and encoder design. These findings suggest the potential of RL-based approaches to understand neural representations of naturalistic faces.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.