Back

Leveraging publicly available datasets and machine learning approaches for predicting the health benefits of fermented foods

McDaniel, E. A.; Edillor, C.; Schertler, M.; Dutton, R. J.

2026-03-05 microbiology
10.64898/2026.03.05.709865 bioRxiv
Show abstract

Fermented foods are an ancient, near universal component of human dietary culture and are increasingly recognized for their health benefits. Bioactive peptides and biosynthetic gene clusters (BGCs) produced by microbes during fermentation have been shown to be key mediators of human health benefits, such as ACE inhibitors and antibacterial bacteriocins. To broadly map this potential, we leveraged the growing abundance of publicly available fermented food datasets alongside recent advances in machine learning models for bioactivity prediction. We collected, curated, and re-analyzed multiple publicly available multi-omics datasets from diverse fermented foods, enabling cross-food comparisons of microbial and molecular profiles. Our analyses include profiling a curated database of [~]1,300 species-representative genomes across hundreds of fermented food metagenomes, predicting genome-encoded BGCs and peptides from 11,500 bacterial genomes, and applying machine learning classification models to predict the bioactivity of thousands of genome-encoded and peptidomics-detected peptides. These models predict 17 different bioactivities, providing novel, testable predictions for downstream experimental characterization. Most importantly, all curated resources, underlying computational tools, and resulting datasets are publicly available to the community, with an emphasis on creating tools and resources that are user-friendly and empower the community to generate more efficient and predictive models for future fermented foods research. A descriptive list of all generated resources is available below, with all computational tools available on GitHub at https://github.com/MicrocosmFoods and all raw files and datasets available on Zenodo at https://zenodo.org/communities/microcosmfoods/. ResourcesO_LIDataset with [~]13,500 genomes from [~]3,000 different samples representing 150 foods and 50 countries C_LIO_LISpecies-representative dataset with [~]1,300 genomes based on 95% ANI, with full functional annotations available as an explorable collection on SeqHub C_LIO_LIStrain-level representative dataset with [~]4,300 genomes based on 99% ANI, available as a free and publicly available narrative on KBase for users to run their own analyses with C_LIO_LIWorkflows for profiling metagenomic samples, performing functional annotation on bacterial genomes, and predicting peptide bioactivity using ML models C_LIO_LIPredicted biosynthetic gene clusters and peptides for [~]11,500 bacterial genomes C_LIO_LIPredicted bioactivities of peptides from [~]11,500 bacterial genomes and 5 proteomics datasets of fermented foods C_LI

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.