Deep Learning Reveals Cross-Modal Neural Representations of Auditory and Visual Mental Imagery in MEG
Schüller, A.; Jehn, C.; Stegmaier, M.; Riegel, J.; Reichenbach, T.
Show abstract
Mental imagery provides a unique window into the brains ability to internally simulate sensory experiences, offering valuable insights for both cognitive neuroscience and brain-computer interface (BCI) research. This study examined the neural representations of imagined auditory and visual stimuli using magnetoen-cephalography (MEG) and assessed the ability of machine learning models to decode these mental processes. MEG data were recorded from 18 right-handed participants during auditory and visual imagery tasks and source-reconstructed within modality-specific cortical regions of interest. We compared a convolutional neural network (CNN) and a linear logistic regression model within a subject-specific classification frame-work. Both approaches achieved above-chance decoding accuracies, with the CNN outperforming the linear model in the auditory task, whereas the linear model showed slightly higher accuracy for visual imagery. Notably, the CNN achieved significant decoding performance even when trained on non-task-relevant cortical regions, indicating that imagined stimuli are represented in distributed and partially overlapping neural networks across modalities. This cross-modal decoding capability highlights the potential of deep learning models to capture complex, multimodal neural patterns and suggests that future brain-computer interfaces could benefit from integrating auditory and visual information. A secondary, behavioral analysis revealed correlation of memory capacity and individual learning preferences with decoding performances, suggesting that individual cognitive differences may further shape the quality of neural representations. Together, these findings advance our understanding of cross-modal mental imagery and point toward more flexible and personalized approaches in BCI design. New and NoteworthyBy comparing linear and deep classifiers, this work shows that convolutional networks capture rich, cross-modal neural representations of auditory and visual mental imagery in MEG. Significant decoding from non-task-relevant regions indicates distributed cortical engagement, highlighting deep learnings potential for robust, modality-independent brain-computer interfaces.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rapid invisible frequency tagging reveals nonlinear integration of auditory and visual information 96%
- Alpha oscillations do not implement gain control in early visual cortex but rather gating in parieto-occipital regions 96%
- Selective Attention Modulates Neural Envelope Tracking of Informationally Masked Speech in Healthy Older Adults 95%
Similar papers in this journal
Similar papers in this journal
- Neural speech tracking benefit of lip movements predicts behavioral deterioration when the speaker's mouth is occluded 97%
- No evidence of musical training influencing the cortical contribution to the speech-FFR and its modulation through selective attention 96%
- Eye Movements in Silent Visual Speech track Unheard Acoustic Signals and Relate to Hearing Experience 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.