Reverse Double-Dipping: When Data Dips You, Twice--Stimulus-Driven Information Leakage in Naturalistic Neuroimaging
Kim, S.-G.
Show abstract
This article elucidates a methodological pitfall of cross-validation for evaluating predictive models applied to naturalistic neuroimaging data--namely, reverse double-dipping (RDD). In a broader context, this problem is also known as leakage in training examples, which is difficult to detect in practice. RDD can occur when predictive modeling is applied to data from a conventional neuroscientific design, characterized by a limited set of stimuli repeated across trials and/or participants. It results in spurious predictive performances due to overfitting to repeated signals, even in the presence of independent noise. Through comprehensive simulations and real-world examples following theoretical formulation, the article underscores how such information leakage can occur and how severely it could compromise the results and conclusions when it is combined with widely spread informal reverse inference. The article concludes with practical recommendations for researchers to avoid RDD in their experiment design and analysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Flexible Multi-Step Hypothesis Testing of Human ECoG Data using Cluster-based Permutation Tests with GLMEs 97%
- The relationship between frequency content and representational dynamics in the decoding of neurophysiological data 96%
- Post-hoc modification of linear models: combining machine learning with domain information to make solid inferences from noisy data 96%
Similar papers in this journal
- Representational similarity learning reveals a graded multi-dimensional semantic space in the human anterior temporal cortex 96%
- The Individualized Neural Tuning Model: Precise and generalizable cartography of functional architecture in individual brains 96%
- Dynamical models reveal anatomically reliable attractor landscapes embedded in resting state brain networks 96%
Similar papers in this journal
- Decoding continuous variables from event-related potential (ERP) data with linear support vector regression (SVR) using the Decision Decoding Toolbox (DDTBOX) 96%
- The Neural Response at the Fundamental Frequency of Speech is Modulated by Word-level Acoustic and Linguistic Information 95%
- Correcting for Superficial Bias in 7T Gradient Echo fMRI 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.