Data-driven estimates of the reproducibility of univariate BWAS are biased.
Burns, C. D. G.; Fracasso, A.; Rousselet, G. A.
Show abstract
Recent studies have used big neuroimaging datasets to answer an important question: how many subjects are required for reproducible brain-wide association studies? These data-driven approaches could be considered a framework for testing the reproducibility of several neuroimaging models and measures. Here we test part of this framework, namely estimates of statistical errors of univariate brain-behaviour associations obtained from resampling large datasets with replacement. We demonstrate that reported estimates of statistical errors are largely a consequence of bias introduced by random effects when sampling with replacement close to the full sample size. We show that future meta-analyses can largely avoid these biases by only resampling up to 10% of the full sample size. We discuss implications that reproducing mass-univariate association studies requires tens-of-thousands of participants, urging researchers to adopt other methodological approaches.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- No effect of additional education on long-term brain structure: a preregistered natural experiment in thousands of individuals. 96%
- Intrinsic dynamic shapes responses to external stimulation in the human brain 95%
- Reconfigurations of cortical manifold structure during reward-based motor learning 95%
Similar papers in this journal
- The connectome spectrum as a canonical basis for a sparse representation of fast brain activity 95%
- Evaluating the sensitivity of functional connectivity measures to motion artifact in resting-state fMRI data 95%
- Disentangling cortical functional connectivity strength and topography reveals divergent roles of genes and environment 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.