Back

The unbiased estimation of r2 between two sets of noisy neural responses

Pospisil, D.; Bair, W. A.

2021-03-30 neuroscience
10.1101/2021.03.29.437413 bioRxiv
Show abstract

The Pearson correlation coefficient squared, r2, is often used in the analysis of neural data to estimate the relationship between neural tuning curves. Yet this metric is biased by trial-to-trial variability: as trial-to-trial variability increases, measured correlation decreases. Major lines of research are confounded by this bias, including the study of invariance of neural tuning across conditions and the similarity of tuning across neurons. To address this, we extend the estimator, [Formula], developed for estimating model-to-neuron correlation to the neuron-to-neuron case. We compare the estimator to a prior method developed by Spearman, commonly used in other fields but widely overlooked in neuroscience, and find that our method has less bias. We then apply our estimator to the study of two forms of invariance and demonstrate how it avoids drastic confounds introduced by trial-to-trial variability. Significant StatementQuantifying the similarity between two sets of averaged neural responses is fundamental to the analysis of neural data. A ubiquitous metric of similarity, the correlation coefficient, is attenuated by trial-to-trial variability that depends on a variety of irrelevant factors. Spearman recognized this problem and proposed corrected methods that have been extended over a century. We show this method has large asymptotic biases and derive a novel estimator to overcome this. Despite the frequent use of the correlation coefficient in neuroscience, consensus on how to address this fundamental statistical issue has not been reached. We both explicate this issue in a neuroscience setting while at the same time making major strides in addressing it.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.