Back

Habitat-specificity in SAR11 is associated with a handful of genes under high selection

Tucker, S. J.; Freel, K. C.; Eren, A. M.; Rappe, M. S.

2024-12-24 microbiology
10.1101/2024.12.23.630198 bioRxiv
Show abstract

The order Pelagibacterales (SAR11) is the most abundant group of heterotrophic bacteria in the global surface ocean, where individual sublineages likely play distinct roles in oceanic biogeochemical cycles. Yet, understanding the determinants of niche partitioning within SAR11 has been a formidable challenge due to the high genetic diversity within individual SAR11 sublineages and the limited availability of high-quality genomes from both cultivation and metagenomic reconstruction. Here, we take advantage of 71 new SAR11 genomes from strains we isolated from the tropical Pacific Ocean to evaluate the distribution of metabolic traits across the Pelagibacteraceae, a recently classified family within the order Pelagibacterales encompassing subgroups Ia and Ib. Our analyses of metagenomes generated from stations where the strains were isolated reveals distinct habitat preferences across SAR11 genera for coastal or offshore environments, and subtle but systematic differences in metabolic potential that support these observations. We also observe higher levels of selective forces acting on habitat-specific metabolic genes linked to SAR11 fitness and polyphyletic distributions of habitat preferences and metabolic traits across SAR11 genera, suggesting that contrasting lifestyles have emerged across multiple lineages independently. Together, these insights reveal niche-partitioning within sympatric and parapatric populations of SAR11 and demonstrate that the immense genomic diversity of SAR11 bacteria naturally segregates into ecologically and genetically cohesive units, or ecotypes, that vary in spatial distributions in the tropical Pacific.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.