Back

Genomic insights into Lactobacillaceae: Analyzing the Alleleome of core pangenomes for enhanced understanding of strain diversity and revealing Phylogroup-specific unique variants

Harke, A. S.; Josephs-Spaulding, J.; Mohite, O. S.; Chauhan, S. M.; Ardalani, O.; Palsson, B.; Phaneuf, P. V.

2023-09-22 genomics
10.1101/2023.09.22.558971 bioRxiv
Show abstract

The Lactobacillaceae familys significance in food and health, combined with available strain-specific genomes, enables genome assessment through pangenome analysis. The Alleleome of the core pangenomes of the Lactobacillaceae family, which identifies natural sequence variations, was reconstructed from the amino acid and nucleotide sequences of the core genes across 2,447 strains of 26 species. It comprised 3.71 million amino acid variants in 29,448 core genes across the family. The alleleome analysis of the Lactobacillaceae family revealed key findings: 1) In the core pangenome, amino acid substitutions prevailed over rare insertions and deletions, 2) Purifying negative selection primarily influenced core gene variations in the family, with diversifying selection noted in L. helveticus. L. plantarums core alleleome was investigated due to its industrial importance. In L. plantarum, the defining characteristics of its core alleleome included: 1) It is highly conserved; 2) Among 235 isolation sources, the primary categories displaying variant prevalence were fermented food, feces, and unidentified sources; 3) It is predominantly characterized by conservative and moderately conservative mutations; and 4) Phylogroup-specific core variant gene analysis identified unique variants (DltX, FabZ1, Pts23B, CspP) in phylogroups I and B which could be used as identifier or validation markers of strain or phylogroup.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.