Gene-language models are whole genome representation learners
Naidenov, B.; Chen, C.
Show abstract
The language of genetic code embodies a complex grammar and rich syntax of interacting molecular elements. Recent advances in self-supervision and feature learning suggest that statistical learning techniques can identify high-quality quantitative representations from inherent semantic structure. We present a gene-based language model that generates whole-genome vector representations from a population of 16 disease-causing bacterial species by leveraging natural contrastive characteristics between individuals. To achieve this, we developed a set-based learning objective, AB learning, that compares the annotated gene content of two population subsets for use in optimization. Using this foundational objective, we trained a Transformer model to backpropagate information into dense genome vector representations. The resulting bacterial representations, or embeddings, captured important population structure characteristics, like delineations across serotypes and host specificity preferences. Their vector quantities encoded the relevant functional information necessary to achieve state-of-the-art genomic supervised prediction accuracy in 11 out of 12 antibiotic resistance phenotypes. TeaserDeep transformers capture and encode gene language content to derive versatile latent embeddings of microbial genomes.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpreting tree ensemble machine learning models with endoR 95%
- DeepDynaForecast: Phylogenetic-informed graph deep learning for epidemic transmission dynamic prediction 95%
- Decoding the Language of Microbiomes: Leveraging Patterns in 16S Public Data using Word-Embedding Techniques and Applications in Inflammatory Bowel Disease 94%
Similar papers in this journal
- Rapid discovery of novel prophages using biological feature engineering and machine learning 95%
- Integrated de novo Gene Prediction and Peptide Assembly of Metagenomic Sequencing Data 94%
- Discovering Governing Equations of Biological Systems through Representation Learning and Sparse Model Discovery 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.