Back

LactoTypeDB: a regenerable, type-anchored 16S rRNA gene reference for species-level identification of the Lactobacillaceae in foods

Oliphant, S. A.; Gardner, J. M.; Jiranek, V.; Sumby, K. M.

2026-08-14 bioinformatics
10.64898/2026.08.12.744342 bioRxiv
Show abstract

Amplicon surveys of fermented and spoiled foods routinely resolve Lactobacillaceae, the lactic acid bacteria responsible for many food and beverage fermentations, only to genus, whereas registers such as the Inventory of Microbial Food Cultures require species-level identification. This shortfall arises from the 16S rRNA genes limited, region-dependent resolution and from incomplete, non-type-strain-anchored references that silently reassign missing species to their nearest relative. We built LactoTypeDB, a regenerable, type-anchored reference covering 434 of the familys 441 species and all 37 genera and substituted it into the Living Tree Project release LTP 08_2023 the fields default classifier uses. This eliminated species-level misassignment of type strains in all regions tested and cut misassignment of 10,329 other sequences from the same species from 1,374 errors down to 3 when the full-length 16S rRNA gene was used. Applied unmodified to 11,612 V3-V4 distinct sequences from a published survey of two meat production lines, the workflow returned a species for 213 and a genus for 5,926, and flagged 3,495 as undescribed candidates, more than a third of them nearest to Dellaglioa, a genus that includes a meat-spoilage organism tracked in that survey. The ambiguity that remains is the markers, since V3-V4 collapses 417 of the 434 species into 27 groups it cannot separate. For food microbiology laboratories, the practical change is that a species call from this family can now be trusted where the marker allows it, and a sequence matching nothing becomes a candidate worth isolating rather than a limitation to work around.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.