The MTIST platform: a microbiome time series inference standardized test simulation, dataset, and scoring systems
Hussey, G. A.; Zhang, C.; Sullivan, A. P.; Fenyo, D.; Schluter, J.
Show abstract
The human gut microbiome is promising therapeutic target, but development of interventions is hampered by limited understanding of the microbial ecosystem. Therefore, recent years have seen a surge in the engineering of inference algorithms seeking to unravel rules of ecological interactions from metagenomic data. Research groups score algorithmic performance in a variety of different ways, however, there exists no unified framework to score and rank each inference approach. The machine learning field presents a useful solution to this issue: a unified set of validation data and accompanying scoring metric. Here, we present MTIST: a platform for benchmarking microbial ecosystem inference tools. We use a generalized Lotka-Volterra framework to simulate microbial abundances over time, akin to what would be obtained by quantitative metagenomic sequencing studies or lab experiments, to generate a massive in silico training dataset (MTIST) for algorithmic validation, as well as an "ecological sign" score (ES score) to rate them. MTIST comprises 24,570 time series of microbial abundance data packaged into 648 datasets. Together, the MTIST dataset and the ES score serve as a platform to develop and compare microbiome ecosystem inference approaches.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BEEM-Static: Accurate inference of ecological interactions from cross-sectional metagenomic data 98%
- Coarse-Grained Model of Serial Dilution Dynamics in Synthetic Human Gut Microbiome 95%
- Linear-regression-based algorithms can succeed at identifying microbial functional groups despite the nonlinearity of ecological function 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.