Back

SimMiL: Simulating Microbiome Longitudinal Data

Weaver, N. E.; Hendricks, A.

2024-03-20 bioinformatics
10.1101/2024.03.18.585571 bioRxiv
Show abstract

0.Structured AbstractO_ST_ABSMotivationC_ST_ABSThe quantity of statistical tools designed for omics data analysis has grown rapidly with the ability to collect large sets of human health data, particularly longitudinal data sets. Most tools are assessed for performance using simulated datasets constructed to mimic a handful of relevant characteristics from real world data sets. Consequently, the simulated data sets, and their respective simulation frameworks, are too narrow in scope to qualify as a standard for assessment in longitudinal omics analyses. ResultsHere we present the flexible and accessible simulation framework and software package called SimMiL (Simulating Microbiome Longitudinal data) capturing three general components of longitudinal microbiome data: (i) absence/presence of microbes, (ii) individual microbe abundance, and (iii) microbiome community composition over time. The framework is assessed by replicating the Type I error and Power analyses of a broad range of statistical tools (MirKAT, repeated measures permANOVA, and a modified kernel association test). Software AvaliabilityThe simulation framework is at https://github.com/nweaver111/SimMiL

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.