An evaluation of reproducibility and errors in published sample size calculations performed using G*Power
Thibault, R. T.; Zavalis, E. A.; Malicki, M.; Pedder, H.
Show abstract
BackgroundPublished studies in the life and health sciences often employ sample sizes that are too small to detect realistic effect sizes. This shortcoming increases the rate of false positives and false negatives, giving rise to a potentially misleading scientific record. To address this shortcoming, many researchers now use point-and-click software to run sample size calculations. ObjectiveWe aimed to (1) estimate how many published articles report using the G*Power sample size calculation software; (2) assess whether these calculations are reproducible and (3) error-free; and (4) assess how often these calculations use G*Powers default option for mixed-design ANOVAs--which can be misleading and output sample sizes that are too small for a researchers intended purpose. MethodWe randomly sampled open access articles from PubMed Central published between 2017 and 2022 and used a coding form to manually assess 95 sample size calculations for reproducibility and errors. ResultsWe estimate that more than 48,000 articles published between 2017 and 2022 and indexed in PubMed Central or PubMed report using G*Power (i.e., 0.65% [95% CI: 0.62% - 0.67%] of articles). We could reproduce 2% (2/95) of the sample size calculations without making any assumptions, and likely reproduce another 28% (27/95) after making assumptions. Many calculations were not reported transparently enough to assess whether an error was present (75%; 71/95) or whether the sample size calculation was for a statistical test that appeared in the results section of the publication (48%; 46/95). Few articles that performed a calculation for a mixed-design ANOVA unambiguously selected the non-default option (8%; 3/36). ConclusionPublished sample size calculations that use G*Power are not transparently reported and may not be well-informed. Given the popularity of software packages like G*Power, they present an intervention point to increase the prevalence of informative sample size calculations.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 95%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 94%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 94%
Similar papers in this journal
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 95%
- Evaluation of statistical methods used to meta-analyse results from interrupted time series studies: a simulation study 94%
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 94%
Similar papers in this journal
- Data-driven Prior Elicitation for Bayes Factors in Cox Regression for Nine Subfields in Biomedicine 94%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 93%
- Modelling the impact of behavioural interventions during pandemics: A systematic review 93%
Similar papers in this journal
- Does advance contact with research participants increase response to questionnaires: A Systematic Review and meta-Analysis 94%
- Does pre-notification increase questionnaire response rates: a nested randomised control trial 94%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.