PHLAME: A benchmark for continuous evaluation of host phenotype prediction from shotgun metagenomic data
Barak, N.; Bhattacharya, H.; Asnicar, F.; Sung, J.; Segata, N.; Yassour, M.
Show abstract
BackgroundPredicting host phenotypes from shotgun metagenomic data is essential for translating microbiome research into clinical practice. Despite the development of numerous computational tools for this task, researchers often default to traditional machine learning methods such as Random Forest. This hesitancy to adopt newer methods stems from their complexity as well as the lack of standardized evaluations, as most tools are assessed on different datasets and compared against a limited set of methods. ResultsHere, we introduce LAMPP, a standardized benchmark for evaluating host phenotype prediction methods using gut metagenomic data. LAMPP features a diverse range of prediction tasks and enables consistent, comparative assessments across prediction tools and is available for ongoing benchmarking at https://lampp.yassourlab.com/ ConclusionsOur systematic evaluation of existing tools shows that classic machine learning methods (e.g., Random Forest) perform competitively, offering both ease of use and state-of-the-art results. At the same time, it demonstrates that microbiome-based phenotype prediction remains a challenging problem. By providing a consistent platform for ongoing evaluation, LAMPP motivates the development of innovative tools that perform beyond the current state of the art.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MDITRE: scalable and interpretable machine learning for predicting host status from temporal microbiome dynamics 96%
- parafac4microbiome: Exploratory analysis of longitudinal microbiome data using Parallel Factor Analysis 94%
- Metagenome-assembled genomes of Estonian Microbiome cohort reveal novel species and their links with prevalent diseases 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.