Learning gene interactions and functional landscapes from entire bacterial proteomes
Sethi, P.; Chevrette, M.; Zhou, J.
Show abstract
Unraveling complex gene interactions and understanding their functions in the genomes of bacteria will provide critical advancements in fields including bacterial genome evolution, microbiome studies, as well as drug and natural product discovery. This is a challenging problem due to the structural and functional complexity of bacterial genomes, and issues including poor gene annotation in non-model species. Language models (LMs) provide a feasible framework for learning the complex interactions among genes from a large number of publicly available, unannotated bacterial genomes. However, applications of language models have mostly been limited to developing models trained on short genomic sequences. Here, we introduce the first whole bacteria proteome foundation model to our knowledge. Our model was trained on ESM embeddings of tens of thousands of full-size proteomes and can generate contextual embeddings for individual proteins as well as embeddings representing the entire genome. We show that our model captures gene-gene interactions and genomic integrity. We further demonstrate that the learned embeddings can be used to achieve state-of-the-art performances for downstream tasks such as identifying operons, and predicting genotype-phenotype maps.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 95%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 95%
- Statistical prediction of microbial metabolic traits from genomes 95%
Similar papers in this journal
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 95%
- A k-mer-based maximum likelihood method for estimating distances of reads to genomes enables genome-wide phylogenetic placement. 95%
- Cross-species imputation and comparison of single-cell transcriptomic profiles 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.