Back

maniFasta and the AllOralsDB: simplifying construction of comprehensive reference databases for metaproteomics

Handelmann, C.; Miles, A. K.; Ye, Y.; Freire, M.; Dewhirst, F. E.; Chen, T.; Mark Welch, J.; Kauffman, K. M.

2026-08-10 microbiology
10.64898/2026.08.08.739415 bioRxiv
Show abstract

Metaproteomics aims to capture a taxonomically comprehensive snapshot of proteins in a sample. Design of reference databases is a key aspect of metaproteomic workflows, as these define what is ultimately seen. Databases tailored to focal biomes offer optimal performance, yet their construction often requires drawing on heterogeneous data sources, posing a challenge to reproducibility and documentation. Here we present maniFasta, a tool enabling users to generate standardized, reproducible, and robustly documented protein reference sets from diverse input sources and datatypes. Users provide information on their desired input types and sources, and the output is an integrated database comprising a protein sequence file (FASTA), with harmonized identifiers and standardized headers, and an associated provenance metadata table (manifest). We highlight the value of maniFasta in the context of salivary metaproteomics, addressing the need for a taxonomically comprehensive reference database. The AllOralsDB resource includes human proteins, as well as proteins from bacteria and archaea, fungi and other microeukaryotes, viruses and viroid-like elements, dietary sources, and common contaminants. Together, this work provides a community resource for oral and salivary metaproteomics (https://www.homd.org/ftp/AllOralsDB/), and a versatile and accessible tool for constructing protein databases for metaproteomics generally (https://github.com/KauffmanLab/maniFasta).

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.