Back

Trends in the Distribution of P Values in Epidemiology Journals: Decreased P Hacking or Increased Power?

Ackley, S. F.; Andrews, R. M.; Seaman, C.; Flanders, M.; Chen, R.; Wang, J.; Lopes, G.; Sims, K.; Buto, P.; Ferguson, E. L.; Allen, I. E.; Glymour, M.

2025-02-14 epidemiology
10.1101/2025.02.13.25322235 medRxiv
Show abstract

Epidemiologists have advocated for reporting confidence intervals and deemphasizing P values to address long-standing concerns about null-hypothesis statistical-significance testing, P hacking, and reproducibility. It is unknown if efforts to reduce reliance on P values have altered the distribution of P values. For 21,377 abstracts published 2000-2024 in 4 major epidemiology journals, two-sided P values (N=25,360) were calculated from estimates and confidence intervals scraped using ChatGPTs 4o model. We evaluated trends over time to determine whether the empirical distribution of P values changed. We fitted to expected P value distributions and simulated these distributions with and without assuming changes in statistical power over time. Average P values decreased from 2000 to 2024, as did the fraction just below the 0.05 threshold. Fits to models indicate that statistical power increased. Increasing power would reduce average P value while also decreasing P values near the 0.05 threshold--precisely the trends observed in epidemiology journals. Although the frequency of P values near the 0.05 threshold has declined modestly, this likely reflects increases in statistical power rather than decreases in P hacking.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.