Sample pooling inflates error rates in between-sample comparisons: an empirical investigation of the statistical properties of count-based data
Taylor, M.; Vega, N.
Show abstract
Heterogeneity is ubiquitous across individuals in biological data, and sample batching, a form of biological averaging, inevitably loses information about this heterogeneity. The consequences for inference from biologically averaged data are frequently opaque, particularly when the underlying populations are non-normal. Here we investigate a case where biological averaging is common - count-based measurement of bacterial load in individual Caenorhabditis elegans - to empirically determine the consequences of batching. We find that both central measures and measures of variation on individual-based data contain biologically relevant information that is useful for distinguishing between groups, and that batch-based inference readily produces both false positive and false negative results in these comparisons. These results support the use of individual rather than batched samples when possible, illustrate the importance of understanding distributions across individuals within a sample frame, and indicate the need to consider effect size when drawing conclusions from biologically averaged data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.