Multimodal Human Perception of Object Dimensions: Evidence from Deep Neural Networks And Large Language Models
Burger, F.; Varlet, M.; Quek, G.; Grootswagers, T.
Show abstract
Object recognition in the human visual system is implemented within a hierarchy characterised by increasing feature complexity. Here, we investigated whether human-derived dimensions of object knowledge show a similar progressive emergence across layers in deep neural networks (DNNs), and how this emergence is shaped by architecture, learning objective, and stimulus statistics. To test this, we predicted human-derived dimensions from layer-wise activations of multiple DNNs and transformer models trained on large-scale datasets. Results showed that trained DNNs exhibit emergence profiles resembling theoretical expectations from human vision, with behaviourally relevant object dimensions largely absent in early layers, strengthening across layers, and peaking in later layers. Architectural mechanisms such as recurrence and skip connections amplified this encoding, learning objectives redistributed information across layers, and changes in stimulus statistics confirm that hierarchical emergence is a general principle extending to material perception. These findings demonstrate that the hierarchical emergence of human-derived dimensions is a fundamental property of trained networks and highlight design and input factors that shape layer-wise representational organisation, providing hypotheses for the structure of visual representations in the brain.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gain, not concomitant changes in spatial receptive field properties, improves task performance in a neural network attention model 97%
- Unsupervised changes in core object recognition behavior are predicted by unsupervised neural plasticity in inferior temporal cortex 95%
- Invariant neural subspaces maintained by feedback modulation 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.