Back

Upstrap for estimating power and sample size in complex models

Karas, M.; Crainiceanu, C.

2021-08-23 ecology
10.1101/2021.08.21.457220 bioRxiv
Show abstract

Power and sample size calculation are major components of statistical analyses. The upstrap resampling method introduced by Crainiceanu and Crainiceanu (2018) was proposed as a general solution to this problem but has not been assessed in numerical experiments. We evaluate the power estimation properties of the upstrap for target data sets that are larger or smaller than the original observed data set. We also expand the scope of upstrap and propose a solution to estimate the power to detect: (1) an effect size observed in the original data; and (2) an effect size chosen by a researcher. Simulations include the following scenarios: one- and two-sample t-tests; linear regression with both Gaussian and binary outcomes; multilevel mixed effects models with both Gaussian and binary outcomes. In addition, our simulations consider cases where the distribution of a covariate in the target setting is preserved and when it is purposefully changed compared to the original data set. We illustrate the approach using a reanalysis of a cluster-randomized controlled trial of malaria transmission. The GitHub repository with R code used in manuscript analyses is available at https://git.io/J0TH1. The accompanying data are publicly available.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.