Disentangling Superpositions: Interpretable Brain Encoding Model with Sparse Concept Atoms
Zeng, A.; Gallant, J.
Show abstract
Encoding models using word embeddings or artificial neural network (ANN) features reliably predict brain responses to naturalistic stimuli, yet interpreting these models remains challenging. A central limitation is superposition: distinct semantic features become entangled along correlated directions in dense embeddings when latent features outnumber embedding dimensions. This entanglement renders regression weights non-identifiable--different combinations of semantic directions can produce identical predictions, precluding principled interpretation of voxel selectivity. To address this, we introduce the Sparse Concept Encoding Model, which transforms dense embeddings into a higher-dimensional, sparse, non-negative space of learned concept atoms. This transformation yields an axis-aligned semantic basis where each dimension corresponds to an interpretable concept, enabling direct readout of conceptual selectivity from voxel weights. When applied to fMRI data collected during story listening, our model matches the prediction performance of conventional dense models while substantially enhancing interpretability. It enables novel neuroscientific analyses such as disentangling overlapping cortical representations of time, space, and number, and revealing structured similarity among distributed conceptual maps. This framework offers a scalable and interpretable bridge between ANN-derived features and human conceptual representations in the brain.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Model metamers illuminate divergences between biological and artificial neural networks 94%
- Edge-centric functional network representations of human cerebral cortex reveal overlapping system-level architecture 94%
- Generalization in Sensorimotor Networks Configured with Natural Language Instructions 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Emergence of brain-like mirror-symmetric viewpoint tuning in convolutional neural networks 95%
- Different computations over the same inputs produce selective behavior in algorithmic brain networks 94%
- A conserved code for anatomy: Neurons throughout the brain embed robust signatures of their anatomical location into spike trains. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.