Back

Reverse Double-Dipping: When Data Dips You, Twice--Stimulus-Driven Information Leakage in Naturalistic Neuroimaging

Kim, S.-G.

2025-04-05 neuroscience
10.1101/2025.04.01.646146 bioRxiv
Show abstract

This article elucidates a methodological pitfall of cross-validation for evaluating predictive models applied to naturalistic neuroimaging data--namely, reverse double-dipping (RDD). In a broader context, this problem is also known as leakage in training examples, which is difficult to detect in practice. RDD can occur when predictive modeling is applied to data from a conventional neuroscientific design, characterized by a limited set of stimuli repeated across trials and/or participants. It results in spurious predictive performances due to overfitting to repeated signals, even in the presence of independent noise. Through comprehensive simulations and real-world examples following theoretical formulation, the article underscores how such information leakage can occur and how severely it could compromise the results and conclusions when it is combined with widely spread informal reverse inference. The article concludes with practical recommendations for researchers to avoid RDD in their experiment design and analysis.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.