Back

Systematic comparison of color representations between humans and deep neural networks: towards predicting human color perception in a vast color space

Wickramanayaka, N. R.; Oizumi, M.

2025-12-13 neuroscience
10.64898/2025.12.10.693611 bioRxiv
Show abstract

The representational structure of human color perception, particularly within vast color spaces, remains incompletely understood. While representational structures for dozens of colors have been studied, exploring structures spanning hundreds or thousands of colors has been infeasible due to the time costs of psychophysical experiments. Given these constraints, deep neural networks (DNNs) have attracted attention as a promising tool for providing proxies or predictions of human perception beyond the scope of psychophysical experiments. However, it remains unclear which DNN models possess internal representations that geometrically align with human color perception. Furthermore, it is unclear which learning paradigm enables DNNs to acquire a color representation that is geometrically aligned with that of humans. Here, we systematically investigate which learning paradigm enables DNNs to produce a color representation that is structurally congruent with that of humans, with a focus on three types: self-supervised learning (SSL) that trains on images alone, supervised learning (SL) that trains on images with category labels, and contrastive language-image pre-training (CLIP) that trains on image-text pairs. We compared internal representations of DNNs with the human similarity judgments of 93 colors using a rigorous unsupervised method termed Gromov-Wasserstein Optimal Transport (GWOT), which reveals whether the representational structures of humans and models align at the fine-item level. Our results show that only the CLIP paradigm acquires color representations that strongly align with human data at the fine-item level. Furthermore, when we leveraged a key advantage of DNNs and investigated a substantial representational structure of 4096 colors, the human-aligned CLIP models consistently converged on a non-trivial distorted ring-like structure, which presents a plausible prediction for the large-scale human color representation. Our work demonstrates an approach for exploring unknown territories of human perception through the use of computational models validated in a limited empirical space, and provides predictions for future large-scale psychophysical experiments. Author summaryHow do we perceive the vast world of color? While scientists have mapped how humans organize a few dozen colors, we still do not know the structure of the massive "color map" that might underlie our perception of thousands of colors, as testing this directly is practically impossible. To explore this space, we turned to deep neural networks, a form of AI. Our first step was to identify models that "see" color in a way that matches humans. We compared models trained under three different conditions: on images alone, on images with categorical labels, and on images paired with rich text descriptions. Using a powerful geometric comparison method, we found that only models trained jointly on images and language strongly matched the human color map. This allowed us to use these models as reliable computational proxies. We then used them to do what human experiments currently cannot: chart a vast map of 4,096 colors. The human-aligned models consistently converged on a unique, non-trivial ring-like structure with specific protrusions and indentations. Our work provides the first plausible, testable prediction for the large-scale structure of human color perception and demonstrates a new way to explore otherwise unreachable territories of our perceptual world.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.