Du-IN-v2: Unleashing the Power of Vector Quantization for Decoding Cognitive States from Intracranial Neural Signals
Zheng, H.; Wang, H.-T.; Jiang, W.-B.; Chen, Z.-T.; He, L.; Lin, P.-Y.; Wei, P.-H.; Zhao, G.-G.; Liu, Y.-Z.
Show abstract
While invasive brain-computer interfaces have shown promise for high-performance speech decoding under medical use, the potential of intracranial stereoElectroEn-cephaloGraphy (sEEG), which causes less damage to patients, remains underex-plored. With the rapid progress in representation learning, leveraging abundant pure recordings to further enhance speech decoding becomes increasingly attractive. However, some popular methods pre-train temporal models based on brain-level tokens, overlooking the brains desynchronization nature; others pre-train spatial-temporal models based on channel-level tokens, yet fail to evaluate them on more challenging tasks, e.g., speech decoding, which demands intricate processing in specific brain regions. To tackle these issues, we introduce a general pre-training framework for speech decoding - Du-IN-v2, which can extract contextual embeddings based on region-level tokens through discrete codex-guided mask modeling. To further push its limits, we propose Decoupling Product Quantization (DPQ), where different codexes are designed to extract different parts of brain dynamics. Our model achieves SOTA performance on both the 61-word classification task and the 49-syllable sequence classification task, surpassing all baselines. Model comparison and ablation studies reveal that our design choices, including (i) temporal modeling based on region-level tokens by utilizing 1D depthwise convolution to fuse channels in vSMC and STG regions and (ii) self-supervision by discrete decoupling codex-guided mask modeling, significantly contribute to these performances. Collectively, our approach, inspired by neuroscience findings, capitalizing on region-level representations from specific brain regions, is suitable for invasive brain modeling. It marks a promising neuro-inspired AI approach in BCI.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning neural decoders without labels using multiple data streams 94%
- Direct Speech Reconstruction from Sensorimotor Brain Activity with Optimized Deep Learning Models 93%
- Speech decoding from a small set of spatially segregated minimally invasive intracranial EEG electrodes with a compact and interpretable neural network 93%
Similar papers in this journal
- Harmonizing and aligning M/EEG datasets with covariance-based techniques to enhance predictive regression modeling 94%
- Alignment massive of auditory individual artificial networks with fMRI brain data leads to generalizable improvements in brain encoding and downstream tasks 94%
- NeuroConText: Contrastive Learning for Neuroscience Meta-Analysis with Rich Text Representation 94%
Similar papers in this journal
Similar papers in this journal
- Evidence for transient, uncoupled power and functional connectivity dynamics 94%
- DeepComBat: A Statistically Motivated, Hyperparameter-Robust, Deep Learning Approach to Harmonization of Neuroimaging Data 94%
- 3D-MASNet: 3D Mixed-scale Asymmetric Convolutional Segmentation Network for 6-month-old Infant Brain MR Images 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.