Back

The genetic repertoire of the deep sea: from sequence to structure and function

Guo, Y.; Wang, Z.; Li, D.; Wang, L.; Lan, H.; Guo, F.; Zhao, Z.; Liu, Z.; Meng, L.; Shen, X.; Wang, M.; Zhao, W.; Zhang, W.; Kong, C.; Shi, L.; Sun, Y.; Seim, I.; Jiang, A.; Ma, K.; Su, Z.; Zhang, N.; Ji, Q.; Chen, J.; Chen, K.; Qi, C.; Li, B.; He, B.; Liu, Y.; Zhou, J.; Zheng, Y.; Zhang, H.; Wang, Y.; Han, M.; Yang, T.; Tong, J.; Zhang, Y.; Wang, Z.; Xu, X.; Chen, J.; Liu, Y.; Chen, H.; Zeng, T.; Wei, X.; Li, C.; Yang, H.; Wang, B.; Liu, X.; Shao, C.; Zhang, W.; Gu, Y.; Xiao, X.; Xu, X.; Wang, J.; Mock, T.; Fan, G.; Li, Y.; Liu, S.; Dong, Y.

2026-02-05 bioinformatics
10.64898/2026.02.03.703410 bioRxiv
Show abstract

The deep sea as the largest and maybe most hostile environment on Earth is still underexplored especially regarding its genetic repertoire. Yet, previous work has revealed significant habitat-specific deep-sea biodiversity. Here, we present an integrated deep-sea genetic dataset comprising 502 million nonredundant genes from 2,138 samples and 2.4 million predicted structures, and used it to link specific protein structures with genetic variants associated with life in the deep sea and to assess their biotechnology potential. Combining global sequence analysis with biophysical and biochemical measurements revealed unprecedented sequence diversity, yet substantial structural conservation of proteins. Especially proteins involved in replication, recombination, and repair were identified to be under rapid evolution and with specialized properties. Among these, a structurally divergent helicase exhibited advantages in controlling nanopore sequencing speed. Thus, our work positions the deep sea as a unique evolutionary engine that generates and hosts genetic diversity and bridges genetic knowledge with biotechnology.

Published in Cell Host & Microbe · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.