StickForStats: automated statistical assumption validation for reproducible computational biology
Bharti, V.; Chakraborty, D.
Show abstract
Reproducible computational biology depends on statistical decisions that routine workflows often skip: verifying that a differential-expression tests assumptions hold across all genes, that a strategy-comparison ANOVA is robust to non-normality, or that a meta-analysis is not distorted by publication bias. Surveys consistently find that fewer than 20% of published biomedical studies report checking these assumptions, and existing statistical software leaves validation to the analyst as an optional step. We present StickForStats, an open-source web platform that reframes assumption validation as a default precondition for every analysis. Its Guardian system--a middleware pipeline of eight validators (normality, variance homogeneity, independence, outliers, sample size, modality, linearity, homoscedasticity)--checks assumptions before execution and, on critical violations, reroutes to an appropriate nonparametric alternative with a documented decision trail. At genome scale, applying Guardian to a 91-sample synovial-sarcoma RNA-seq study (GSE271517) cascaded 90.6% of 27,221 genes to a rank-based test and flipped the differential-expression verdict for 553 genes--479 rescued from an under-powered t-test and 74 outlier-driven false positives rejected--materially changing the gene list a biologist would act on. The same automatic validation generalizes across domains: a CRISPR editing-strategy comparison (ANOVA F = 1122, with Guardian recommending Kruskal-Wallis H = 36.6), an ordinal correlation (Pearson r = 0.476 corrected to Spearman {rho} = 0.479), and a sixteen-trial clinical meta-analysis revealing severe publication bias (Eggers t = -5.78, p < 0.001); a complementary module extends the same validators to published manuscripts, checking claims against CONSORT, STROBE, ICH-E9, and JARS-Quant reporting standards. By making assumption validation automatic and transparent, StickForStats targets a tractable, under-served contributor to irreproducibility. The platform is MIT-licensed, validated against SciPy and R, and freely available at https://github.com/visvikbharti/stickforstats_new. Author summaryMost scientific conclusions rest on statistical tests, and every test comes with fine print: assumptions about the data that must hold for the result to be trustworthy. In practice, this fine print is often left unchecked. Surveys find that fewer than one in five published studies reports verifying these assumptions, partly because popular software treats the check as an optional extra that busy researchers easily skip. When the assumptions are ignored, a study can report a difference that is not really there. We built StickForStats to make this checking automatic. Before it runs any statistical test, our platform inspects the data, reports whether each assumption is met, and--if a serious problem is found--switches to a more appropriate method and records why. On four real biomedical datasets, including a gene-editing comparison and a large gene-expression study, we show that this safety net changes which findings are flagged as reliable. A companion tool applies the same checks to finished manuscripts, helping catch reporting problems before publication. By turning assumption checking from something you must remember into something that happens by default, we aim to make everyday analyses more reproducible.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ShinyLearner: A containerized benchmarking tool for machine-learning classification of tabular data 94%
- Reproducible processing of TCGA regulatory networks 94%
- Strategies and Techniques for Quality Control and Semantic Enrichment with Multimodal Data: A Case Study in Colorectal Cancer with eHDPrep 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.