Gaia: A Context-Aware Sequence Search and Discovery Tool for Microbial Proteins
Jha, N.; Kravitz, J.; West-Roberts, J.; Camargo, A.; Roux, S.; Cornman, A.; Hwang, Y.
Show abstract
Protein sequence similarity search is fundamental to genomics research, but current methods are typically not able to consider crucial genomic context information that can be indicative of protein function, especially in microbial systems. Here we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising over 85M protein clusters (defined at 90% sequence identity) from 131,744 microbial genomes. We compare the sequence, structure and context sensitivity of gLM2 embedding-based search against existing tools like MMseqs2 and Foldseek. We showcase Gaia-enabled discoveries of phage tail proteins and siderophore synthesis loci that were previously difficult to annotate with traditional tools. Gaia search is freely available at https://gaia.tatta.bio.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Deciphering the Biosynthetic Potential of Microbial Genomes Using a BGC Language Processing Neural Network Model 96%
- Domainator, a flexible software suite for domain-based annotation and neighborhood analysis, identifies proteins involved in antiviral systems 95%
- Targeted genome mining with GATOR-GC maps the evolutionary landscape of biosynthetic diversity 95%
Similar papers in this journal
Similar papers in this journal
- iPRESTO: automated discovery of biosynthetic sub-clusters linked to specific natural product substructures 96%
- ResMiCo: increasing the quality of metagenome-assembled genomes with deep learning 96%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.