Investigating the temporal dynamics and modelling of mid-level feature representations in humans
Karapetian, A.; Lenders, A.; Bawa, V.; Pflaum, M.; Leuner, R.; Roig, G.; Dwivedi, K.; Cichy, R. M.
Show abstract
Scene perception is a key function of biological visual systems. According to the hierarchical processing view, scene perception in the human brain begins with low-level features, progresses to mid-level features, and ends with high-level features. While low- and high-level feature processing is well-studied, research on mid-level features remains limited. Here, we addressed this gap by investigating when mid-level features are processed in humans using a novel stimulus set of naturalistic scenes as images and videos, accompanied with ground-truth annotations for five mid-level features (reflectance, lighting, world normals, scene depth and skeleton position), and two framing features: one low-level (edges) and one high-level feature (action). To reveal when low-, mid- and high-level features are represented in the brain, we collected electroencephalography (EEG) data from human participants during stimulus presentation and trained encoding models to predict EEG data from ground-truth annotations. We revealed that mid-level features were best represented between [~]100 and [~]250 ms post-stimulus, between low- and high-level features. Moreover, we assessed scene- and action-trained convolutional neural networks (CNNs) as models of mid-level feature processing in humans. We found a comparable processing order for mid-but not low- or high-level features with humans. Overall, our results characterize mid-level feature processing in humans in the temporal domain and reveal CNNs as suitable models of the processing hierarchy of mid-level vision in humans.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EEG decoding reveals neural predictions for naturalistic material behaviors 98%
- Disentangling the Independent Contributions of Visual and Conceptual Features to the Spatiotemporal Dynamics of Scene Categorization 96%
- The relative coding strength of object identity and nonidentity features in human occipito-temporal cortex and convolutional neural networks 96%
Similar papers in this journal
- The representational dynamics of the animal appearance bias in human visual cortex are indicative of fast feedforward processing 96%
- Encoding neural representations of time-continuous stimulus-response transformations in the human brain with advanced deep neural networks 96%
- The Individualized Neural Tuning Model: Precise and generalizable cartography of functional architecture in individual brains 96%
Similar papers in this journal
Similar papers in this journal
- Differential involvement of EEG oscillatory components in sameness vs. spatial-relation visual reasoning tasks 95%
- Sensory and perceptual decisional processes underlying the perception of reverberant auditory environments 94%
- Measuring stimulus-evoked neurophysiological differentiation in distinct populations of neurons in mouse visual cortex 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.