Back

Einstein from Noise: Statistical Analysis

Balanov, A.; Huleihel, W.; Bendory, T.

2024-07-07 molecular biology
10.1101/2024.07.06.602366 bioRxiv
Show abstract

"Einstein from noise" (EfN) is a prominent example of the model bias phenomenon, where systematic errors in the statistical model lead to spurious but consistent estimates. In the EfN experiment, one falsely believes that a set of observations contains noisy, shifted copies of a template signal (e.g., an Einstein image), whereas in reality, it contains only pure noise observations. To estimate the signal, the observations are first aligned with the template using cross-correlation and then averaged. Although the observations contain nothing but noise, it was recognized early on that this process produces a signal that resembles the template signal! This model bias pitfall was at the heart of a central scientific controversy about validation techniques in structural biology. This paper provides a comprehensive statistical analysis of the EfN phenomenon above. We show that the Fourier phases of the EfN estimator (namely, the average of the aligned noise observations) converge to the Fourier phases of the template signal, thereby explaining the observed structural similarity. Additionally, we prove that the convergence rate of Fourier phases is inversely proportional to the number of noise observations and, in the high-dimensional regime, to the Fourier magnitudes of the template signal. Moreover, in the high-dimensional regime, the EfN estimator converges to a scaled version of the template signal. This work not only deepens the theoretical understanding of the EfN phenomenon but also highlights potential pitfalls in template matching techniques and emphasizes the need for careful interpretation of noisy observations across disciplines in engineering, statistics, physics, and biology.

Published in IEEE Transactions on Signal Processing · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Journal of Mathematical Biology
40 papers in training set
Top 0.1%
18.9%
2
Physical Review Research
49 papers in training set
Top 0.1%
15.2%
3
Journal of Theoretical Biology
162 papers in training set
Top 0.3%
8.1%
4
PLOS ONE
5266 papers in training set
Top 24%
6.9%
5
Journal of The Royal Society Interface
235 papers in training set
Top 0.6%
5.6%
50% of probability mass above
6
Statistics in Medicine
40 papers in training set
Top 0.1%
5.0%
7
Entropy
21 papers in training set
Top 0.1%
4.1%
8
Biophysical Journal
631 papers in training set
Top 2%
3.6%
9
PLOS Computational Biology
1863 papers in training set
Top 9%
3.3%
10
The Journal of Chemical Physics
56 papers in training set
Top 0.2%
2.4%
11
Bulletin of Mathematical Biology
92 papers in training set
Top 0.7%
2.2%
12
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 23%
2.2%
13
Scientific Reports
3612 papers in training set
Top 52%
1.8%
14
Biology Methods and Protocols
61 papers in training set
Top 1%
1.5%
15
The Annals of Applied Statistics
19 papers in training set
Top 0.2%
1.1%
16
Optics Express
26 papers in training set
Top 0.2%
1.1%
17
BMC Bioinformatics
457 papers in training set
Top 5%
0.9%
18
MethodsX
16 papers in training set
Top 0.2%
0.9%
19
eLife
5828 papers in training set
Top 64%
0.9%
20
Nature Communications
5641 papers in training set
Top 59%
0.6%
21
Biometrics
23 papers in training set
Top 0.4%
0.6%
22
Physical Review E
112 papers in training set
Top 1%
0.6%