Back

RMeDPower2 for Biology: guiding the design, experimental structure and analyses of experiments generating repeated measures datasets

Shin, M.-G.; Amirani, N.; Lam, S.; Al Bistami, N.; Raja, K.; Vertudes, E.; Kaye, J. A.; Thomas, R.; Finkbeiner, S.

2026-07-22 cell biology
10.1101/2022.07.18.500490 bioRxiv
Show abstract

Lack of experimental reproducibility has plagued efforts to understand biology at both basic biomedical and preclinical levels. The cause is often improperly powered experiments and the use of inadequate statistical tools. To overcome these problems, we developed RMeDPower2, a complete, user-friendly package of tools in R that will allow scientists that are not deeply familiar with statistical analyses to predict the scope and size of biological data they need when conducting experiments with a repeated measures design. RMeDPower2 is based on Generalized Linear Mixed Effects Models (GLMM), which are better suited to the statistical analysis of these experiments than ANOVA or t-tests. We illustrate the use of RMeDPower2 and compare it to t- test for power calculations, using our own pilot studies of iPSC-derived motor neurons (iMNs) from sporadic ALS (sALS) patients versus healthy controls. We report that sALS iMNs display reduced numbers of soma- emanating processes compared to control iMNs using RMeDPower2. We expect RMeDPower2 to find applications far beyond cell assays, from single-cell RNAseq experiments to brain slice electrophysiology or animal behavior. MotivationThe lack of rigor and reproducibility in biomedical research has caused a crisis that has been highlighted in the popular literature and has become a focus for the National Institutes of Health1-3. It has been estimated that the majority of published empirical observations cannot be reproduced4-9, rendering nearly futile any effort to build on these observations to further our understanding of basic biological mechanisms or design effective therapeutic approaches. Further, the resources and time spent attempting to reproduce findings from low-quality or incorrectly acquired data are estimated to cost the global scientific community about 200 billion dollars per year10. The root cause lies in experimental designs that are not structured or powered adequately for conclusive statistical analyses. Since all biomedical researchers cannot be expected to have a deep knowledge of statistics or easy access to trained statisticians, tools are desperately needed to help them check the design of their experiments and apply adequate statistical power estimation. Not only could this improve our confidence in scientific outcomes, it could help make biological experiments more time-efficient and cost-effective. For example, if a researcher could estimate how many experiments should be performed and how many cell lines, animals or tissue samples should be collected to achieve sufficient statistical power to test their hypothesis, they may adjust their experimental design to fit their time or budgetary constraints without jeopardizing the quality of their findings. Another source of scientific errors comes from technologies such as scRNA-seq, whose advances are leading to a rapid increase in studies involving so-called "pseudo-replication", which treats non-independent measures as if they were independent. For example, carrying out multiple measurements on a single sample instead of using separate, independent samples would represent non-independent replication. The risk of pseudo-replication11 (illustrated further below), can be remedied by the implementation of rigorous statistical methods that apply to all aspects of the data arising from such designs.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.