Back

Deeply-sequenced metagenomes and over 1000 draft genomes from the epipelagic to bathypelagic Northeast Pacific Ocean

Jaffe, A. L.; Salcedo, R. S. R.; Lewitton, J. J.; Dekas, A. E.

2026-07-20 microbiology
10.64898/2026.07.18.739368 bioRxiv
Show abstract

The deep ocean water column (>200 meters depth) hosts diverse and active microbial communities that play critical roles in Earths biogeochemical cycles. However, this habitat has been massively undersampled compared to shallower depths where sunlight penetrates, as well as benthic features like hydrothermal vents and seeps. In particular, few existing deep-sea datasets provide adequate spatial resolution to study the impact of physical and geochemical gradients on marine microbial communities over both horizontal and vertical spatial scales. Here, we present a deeply-sequenced, spatially-resolved set of 28 water column metagenomes from a [~]300 km horizontal transect off the coast of central California, spanning 50 to 4000 meters depth. We estimate that the deepest portion of the resulting [~]1.5 terabases of sequencing data comprises nearly 9% of all shotgun metagenomic sequencing currently available from 1000 meters depth or below. Additionally, we generate over one thousand draft genomes describing the bacteria and archaea present in our bulk metagenomic data. These genomes frequently account for the majority of reads sequenced, and show discernable patterns of abundance with depth, underscoring the presence of distinct microbial cohorts inhabiting different ocean compartments. Additionally, the depth of sequencing performed (mean of [~]51.9 gigabases per sample) enables resolution of rare community members that are, on average, more phylogenetically novel than their abundant counterparts, forming an important contribution to genomic catalogues for this habitat. We anticipate that the combined datasets will enable diverse research questions concerning the adaptation, functional potential, and distribution of deep sea microbes as well as their genes and proteins.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.